sk Skill · charlieviettq
agent-evaluation
Evaluate LLM agents and tool-using workflows—task success, tool accuracy, latency/cost, safety, and regression suites. Use when shipping agent features, comparing prompts/models, or debugging agent failures.
Open on skills.sh ↗read 2026-09-17
- installs 8w
- 0
- 30-day movement
- starts with the next reading
- Related entries
- 5
- Connections
- 3
Python
- Host repository
- charlieviettq/awesome-agent-skill
- Allowed tools
- Read, Glob, Grep
- Host stars
- 26
- Host language
- Python