sk Skill · PKU-YuanGroup
bioprobench
Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.
Open on skills.sh ↗read 2026-09-17
- installs 8w
- 0
- 30-day movement
- starts with the next reading
- Related entries
- 1
- Connections
- 0
pythonbibtexPythonmodel-evaluation
- Host repository
- PKU-YuanGroup/OpenAI4S
- Category
- model-evaluation
- Host stars
- 552
- Host language
- Python