BigHugger
sk Skill · agentscope-ai

openjudge

Build custom LLM evaluation pipelines using the OpenJudge framework. Covers selecting and configuring graders (LLM-based, function-based, agentic), running batch evaluations with GradingRunner, combining scores with aggregators, applying evaluation strategies (voting, average), auto-generating graders from data, and analyzing results (pairwise win rates, statistics, validation metrics). Use when the user wants to…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
pythonbashPython
Host repository
agentscope-ai/OpenJudge
Host stars
838
Host language
Python