BigHugger
GH Repository · xlang-ai

OSWorld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

stars
3,148
30-day movement
+52/day
Related entries
60
Connections
1
python/uvagent-appPythonvlmagentbenchmarknatural-language-processingartificial-intelligencelarge-action-modelrpaguireinforcement-learningcode-generationpythonllmlanguage-modelmultimodalcli

OSWorld is a benchmark from xlang-ai for evaluating multimodal agents on open-ended tasks in real computer environments, published at NeurIPS 2024. The repository provides the benchmark itself as a Python toolchain, with topics indicating coverage of both GUI and CLI interaction settings.

You need a standardized, peer-reviewed way to measure how well multimodal agents perform real computer tasks rather than toy environments.

Use it to

  • Evaluate multimodal agents on real computer tasks
  • Compare GUI and CLI agent performance
  • Benchmark LLM and VLM-based agents
  • Support reinforcement-learning and RPA research
  • Reproduce NeurIPS 2024 benchmark results

For Researchers benchmarking multimodal and LLM-based agents

Role
agent-app
Language
Python
Licence
Apache-2.0
Forks
531
Open issues
155
Last push
2026-09-14
Latest release
v0.1.0 · 2024-04-11
topicsbenchmarkmultimodal-agentsguillmevaluationcomputer-use