sk Skill · noah-sheldon
eval-harness
Evaluation harness for LLM and code quality assessment. Covers Pass@k metrics, golden datasets, LLM-as-judge, benchmark suites, regression detection, and eval-driven development workflow.
Open on skills.sh ↗read 2026-09-17
- installs 8w
- 0
- 30-day movement
- starts with the next reading
- Related entries
- 12
- Connections
- 0
yamlpythonjson
- Host repository
- noah-sheldon/ai-dev-kit
- Invocable by
- model
- Host stars
- 13