GH Repository · open-compass
VLMEvalKit
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
- stars
- 4,393
- 30-day movement
- +52/day
- Related entries
- 60
- Connections
- 1
Pythonevaluationvitcliplarge-language-modelsqwenllavaopenaichatgptmulti-modalgeminiopenai-apivqapythonllmgpt-4vgptevalcomputer-visionpytorchclaudegpt4
VLMEvalKit is an open-source Python toolkit for evaluating large multi-modality models, supporting 220+ LMMs and 80+ benchmarks. It is maintained under the OpenCompass project and licensed Apache-2.0.
Use it to benchmark multimodal models across a broad set of standard VQA and vision-language benchmarks from one toolkit.
Use it to
- Evaluate a multimodal model on 80+ benchmarks
- Compare LMMs like GPT-4V, LLaVA, Qwen, and Gemini
- Run visual question answering evaluations
- Benchmark models via OpenAI-compatible APIs
For Researchers and engineers benchmarking vision-language models
- Role
- eval
- Language
- Python
- Licence
- Apache-2.0
- Forks
- 768
- Open issues
- 225
- Last push
- 2026-09-17
- Latest release
- v0.1 · 2024-01-22
topicsevaluationmultimodalvqallmcomputer-visionpytorch