BigHugger
sk Skill · wanshuiyin

experiment-forensics

Audit experiment integrity against the evidence ledger. At L2 (repo + result files present) a fresh cross-model reviewer reads the eval code line-by-line for fake/derived ground truth, score self-normalization, phantom results (a paper number with no backing file/key), dead/uncalled metric code, verified-scope inflation, method-described ≠ method-evaluated drift, synthesized-looking results, placeholder/fake data…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
6
jsonbashPython
Host repository
wanshuiyin/Anti-Autoresearch
Allowed tools
Bash(*), Read, Write, Grep, Glob, mcp__codex__codex
Host stars
154
Host language
Python