BigHugger
sk Skill · kibertoad

benchmark

Run a cat-factory agent benchmark end to end. Use when asked to benchmark, compare, or measure models or prompt versions on the agent tasks (requirement review, code review, implementation) — e.g. "benchmark Llama vs Claude on code review", "compare review@v1 vs a new prompt", "run the benchmark and grade it". Configures the matrix, runs cat-bench, invokes the benchmark-arbiter skill to grade, merges the report, and…

installs 8w
0
30-day movement
starts with the next reading
Related entries
12
Connections
0
bashTypeScript
Host repository
kibertoad/cat-factory
Host stars
8
Host language
TypeScript