BigHugger
sk Skill · agentscope-ai

ref-hallucination-arena

Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP. Measures hallucination rate, per-field accuracy (title/author/year/DOI), discipline breakdown, and year constraint compliance. Supports tool-augmented (ReAct + web search) mode. Use when the user asks to evaluate, benchmark, or compare models on academic reference hallucination, literature…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
yamlpythonjsonbashPython
Host repository
agentscope-ai/OpenJudge
Host stars
838
Host language
Python