BigHugger
sk Skill · khalilbenaz

agent-evaluation-framework

Framework d'évaluation et benchmarking d'agents IA. Métriques, tests, comparaisons et quality assurance. Se déclenche avec "évaluer agent", "agent eval", "benchmark agent", "tester mon agent", "qualité agent", "agent metrics", "agent performance", "LMSYS", "agent accuracy". Also triggers on "evaluate my agent", "agent benchmark", "agent eval harness".

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
yamlpythonHTML
Host repository
khalilbenaz/claude-skills-collection
Host stars
22
Host language
HTML