sk Skill · chandrudp29
llm-evaluator
Evaluate LLM outputs systematically using LLM-as-judge, human evaluation frameworks, and regression testing. Use when assessing model quality, comparing models, or preventing quality regression.
Open on skills.sh ↗read 2026-09-15
- installs 8w
- 0
- 30-day movement
- starts with the next reading
- Related entries
- 1
- Connections
- 0
Pythonregressionllm-as-judgepythonbenchmarkingqualityevaluationllm
- Host repository
- chandrudp29/skillhub
- Version
- 1.0.0
- Host stars
- 13
- Host language
- Python