BigHugger
sk Skill · thiientv

agent-evaluation

Designs and runs reproducible evaluations for AI agents, prompts, tools, skills, and model-backed workflows using realistic datasets, isolated baselines, objective assertions, rubric grading, trajectory analysis, cost/latency tracking, and regression comparison. Use when measuring agent quality, optimizing skill triggering, comparing prompts or models, or gating an AI feature release. Not for ordinary deterministic…

installs 8w
0
30-day movement
starts with the next reading
Related entries
5
Connections
0
Python
Host repository
thiientv/godmode
Host stars
94
Host language
Python