sk Skill · wanshuiyin
experiment-forensics
Audit experiment integrity against the evidence ledger. At L2 (repo + result files present) a fresh cross-model reviewer reads the eval code line-by-line for fake/derived ground truth, score self-normalization, phantom results (a paper number with no backing file/key), dead/uncalled metric code, verified-scope inflation, method-described ≠ method-evaluated drift, synthesized-looking results, placeholder/fake data…
Open on skills.sh ↗read 2026-09-15
- installs 8w
- 0
- 30-day movement
- starts with the next reading
- Related entries
- 1
- Connections
- 6
jsonbashPython
- Host repository
- wanshuiyin/Anti-Autoresearch
- Allowed tools
- Bash(*), Read, Write, Grep, Glob, mcp__codex__codex
- Host stars
- 154
- Host language
- Python