BigHugger
sk Skill · edonadei

evaluate-skill

Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec.

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
1
yamlbashPython
Host repository
edonadei/caliper
Allowed tools
Bash
Host stars
158
Host language
Python