BigHugger
sk Skill · Aperivue

design-ai-benchmarking

Design and validity review for studies that benchmark one or more AI systems against a human-expert panel as the reference. Covers the evaluation question and arm definition, decoupled multi-dimensional rubrics with anchors, planted calibration probes, reviewer-panel construction, inter-rater reliability targets, LLM-as-judge versus human-as-judge adjudication, construct-independence guards, and a structured…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
6
Python
Host repository
Aperivue/medsci-skills
Allowed tools
Read, Write, Edit, Bash, Grep, Glob
Host stars
305
Host language
Python