sk Skill · NoesisVision
nasde-benchmark-calibration
Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: — Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark — Review trial diffs alongside the LLM-as-a-Judge scores in a PR/MR — Investigate why judge scores feel off, too harsh, too lenient, or misaligned with how a human would…
Open on skills.sh ↗read 2026-09-17
- installs 8w
- 0
- 30-day movement
- starts with the next reading
- Related entries
- 1
- Connections
- 0
tomlbashPython
- Host repository
- NoesisVision/nasde-toolkit
- Host stars
- 14
- Host language
- Python