BigHugger
sk Skill · alebgl77

agent-evaluation

Measures a skill's practical value with fixed business tasks, paired with-skill and without-skill trials, explicit scoring, and honest uncertainty and cost reporting. Use when the user says "does this skill help", "benchmark our agent workflow", "compare this skill to the baseline", or "prove the new workflow works".

installs 8w
0
30-day movement
starts with the next reading
Related entries
5
Connections
0
markdownPython
Host repository
alebgl77/claude-inc
Host stars
15
Host language
Python