GH Repository · xlang-ai
OSWorld
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- stars
- 3,148
- 30-day movement
- +52/day
- Related entries
- 60
- Connections
- 1
python/uvagent-appPythonvlmagentbenchmarknatural-language-processingartificial-intelligencelarge-action-modelrpaguireinforcement-learningcode-generationpythonllmlanguage-modelmultimodalcli
OSWorld is a benchmark from xlang-ai for evaluating multimodal agents on open-ended tasks in real computer environments, published at NeurIPS 2024. The repository provides the benchmark itself as a Python toolchain, with topics indicating coverage of both GUI and CLI interaction settings.
You need a standardized, peer-reviewed way to measure how well multimodal agents perform real computer tasks rather than toy environments.
Use it to
- Evaluate multimodal agents on real computer tasks
- Compare GUI and CLI agent performance
- Benchmark LLM and VLM-based agents
- Support reinforcement-learning and RPA research
- Reproduce NeurIPS 2024 benchmark results
For Researchers benchmarking multimodal and LLM-based agents
- Role
- agent-app
- Language
- Python
- Licence
- Apache-2.0
- Forks
- 531
- Open issues
- 155
- Last push
- 2026-09-14
- Latest release
- v0.1.0 · 2024-04-11
topicsbenchmarkmultimodal-agentsguillmevaluationcomputer-use