BigHugger
sk Skill · chandrudp29

llm-evaluator

Evaluate LLM outputs systematically using LLM-as-judge, human evaluation frameworks, and regression testing. Use when assessing model quality, comparing models, or preventing quality regression.

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
Pythonregressionllm-as-judgepythonbenchmarkingqualityevaluationllm
Host repository
chandrudp29/skillhub
Version
1.0.0
Host stars
13
Host language
Python