BigHugger
sk Skill · charlieviettq

agent-evaluation

Evaluate LLM agents and tool-using workflows—task success, tool accuracy, latency/cost, safety, and regression suites. Use when shipping agent features, comparing prompts/models, or debugging agent failures.

installs 8w
0
30-day movement
starts with the next reading
Related entries
5
Connections
3
Python
Host repository
charlieviettq/awesome-agent-skill
Allowed tools
Read, Glob, Grep
Host stars
26
Host language
Python