BigHugger
sk Skill · ancoleman

evaluating-llms

Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks. Use when testing prompt quality, validating RAG pipelines, measuring safety (hallucinations, bias), or comparing models for production deployment.

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
pythonbashPython
Host repository
ancoleman/ai-design-components
Host stars
522
Host language
Python