BigHugger
sk Skill · noah-sheldon

eval-harness

Evaluation harness for LLM and code quality assessment. Covers Pass@k metrics, golden datasets, LLM-as-judge, benchmark suites, regression detection, and eval-driven development workflow.

installs 8w
0
30-day movement
starts with the next reading
Related entries
12
Connections
0
yamlpythonjson
Host repository
noah-sheldon/ai-dev-kit
Invocable by
model
Host stars
13