BigHugger
sk Skill · zjunlp

experiment-audit

Audit the experimental methodology integrity for a specific claim (Checks A–F: GT provenance, score normalization, result-file existence, dead code, scope, eval-type). Uses cross-model review (external LLM reviewer via llm-chat MCP). The output overall_verdict (PASS/WARN/FAIL) is THIS claim's verdict — i.e., whether the claim's experimental process is methodologically sound. Does NOT judge whether the numbers…

installs 8w
0
30-day movement
starts with the next reading
Related entries
2
Connections
8
jsonmarkdownbashPython
Host repository
zjunlp/Mechanist
Allowed tools
Bash(*), Read, Write, Edit, Grep, Glob, Agent, mcp__llm-chat__chat
Host stars
75
Host language
Python