BigHugger
sk Skill · usecompai

skill-evaluation

Evaluate whether a the reference deployment skill actually works — blind runner + separate judge against a bar, rubric scoring, pass-rate on re-runs, and defect diagnosis. Use when asked to evaluate a skill, eval skills, run a blind eval, score against a rubric, test a skill, audit skill quality, grade a skill, review the skill library, or before promoting a new skill to the company master prompt or the whole swarm.…

installs 8w
0
30-day movement
starts with the next reading
Related entries
2
Connections
0
Python
Host repository
usecompai/compound-operations-model
Host stars
26
Host language
Python