sk Skill · ericrisco
agent-eval
Use when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework. NOT building the agent loop, tools or RAG plumbing (that is building-agents).
Open on skills.sh ↗read 2026-09-17
- installs 8w
- 0
- 30-day movement
- starts with the next reading
- Related entries
- 6
- Connections
- 0
jsonlJavaScriptpythonregression-gatellm-as-judgeaiagentsllmevals
- Host repository
- ericrisco/rsc-harness
- Host stars
- 84
- Host language
- JavaScript