BigHugger
sk Skill · PKU-YuanGroup

bioprobench

Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
pythonbibtexPythonmodel-evaluation
Host repository
PKU-YuanGroup/OpenAI4S
Category
model-evaluation
Host stars
552
Host language
Python