BigHugger
GH Repository · THUDM

AgentBench

A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)

stars
3,735
30-day movement
starts with the next reading
Related entries
60
Connections
1
pythonagent-frameworkPythonllm-agentllmgpt-4chatgpt

AgentBench is a benchmark suite for evaluating large language models as agents, published at ICLR'24. It is a Python toolchain from THUDM under an Apache-2.0 license.

Use it to systematically measure how well LLMs perform in agentic settings rather than relying on ad-hoc tests.

Use it to

  • Benchmark LLMs on agentic tasks
  • Compare models like GPT-4 and ChatGPT
  • Reproduce published ICLR'24 agent evaluations
  • Extend the benchmark with Python tooling

For Researchers and engineers evaluating LLM agents

Role
agent-framework
Language
Python
Licence
Apache-2.0
Forks
279
Open issues
65
Last push
2026-02-08
topicsbenchmarkevaluationllmllm-agentagentspython