GH Repository · THUDM
AgentBench
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
- stars
- 3,735
- 30-day movement
- starts with the next reading
- Related entries
- 60
- Connections
- 1
pythonagent-frameworkPythonllm-agentllmgpt-4chatgpt
AgentBench is a benchmark suite for evaluating large language models as agents, published at ICLR'24. It is a Python toolchain from THUDM under an Apache-2.0 license.
Use it to systematically measure how well LLMs perform in agentic settings rather than relying on ad-hoc tests.
Use it to
- Benchmark LLMs on agentic tasks
- Compare models like GPT-4 and ChatGPT
- Reproduce published ICLR'24 agent evaluations
- Extend the benchmark with Python tooling
For Researchers and engineers evaluating LLM agents
- Role
- agent-framework
- Language
- Python
- Licence
- Apache-2.0
- Forks
- 279
- Open issues
- 65
- Last push
- 2026-02-08
topicsbenchmarkevaluationllmllm-agentagentspython