BigHugger
sk Skill · NoesisVision

nasde-benchmark-calibration

Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: — Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark — Review trial diffs alongside the LLM-as-a-Judge scores in a PR/MR — Investigate why judge scores feel off, too harsh, too lenient, or misaligned with how a human would…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
tomlbashPython
Host repository
NoesisVision/nasde-toolkit
Host stars
14
Host language
Python