BigHugger
sk Skill · shyftlabs

continuum-evaluation

Evaluate agent quality with the EvaluatorAgent, generate golden datasets from a corpus, and run DeepEval/RAGAS metrics over conversations. Invoke when the user asks "test agent quality", "evaluate output", "RAG metrics", "DeepEval", "RAGAS", or "regression-test my agent".

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
pythonbash
Host repository
shyftlabs/continuum
Host stars
84