BigHugger
sk Skill · ericrisco

agent-eval

Use when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework. NOT building the agent loop, tools or RAG plumbing (that is building-agents).

installs 8w
0
30-day movement
starts with the next reading
Related entries
6
Connections
0
jsonlJavaScriptpythonregression-gatellm-as-judgeaiagentsllmevals
Host repository
ericrisco/rsc-harness
Host stars
84
Host language
JavaScript