BigHugger
GH Repository · open-compass

VLMEvalKit

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

stars
4,393
30-day movement
+52/day
Related entries
60
Connections
1
Pythonevaluationvitcliplarge-language-modelsqwenllavaopenaichatgptmulti-modalgeminiopenai-apivqapythonllmgpt-4vgptevalcomputer-visionpytorchclaudegpt4

VLMEvalKit is an open-source Python toolkit for evaluating large multi-modality models, supporting 220+ LMMs and 80+ benchmarks. It is maintained under the OpenCompass project and licensed Apache-2.0.

Use it to benchmark multimodal models across a broad set of standard VQA and vision-language benchmarks from one toolkit.

Use it to

  • Evaluate a multimodal model on 80+ benchmarks
  • Compare LMMs like GPT-4V, LLaVA, Qwen, and Gemini
  • Run visual question answering evaluations
  • Benchmark models via OpenAI-compatible APIs

For Researchers and engineers benchmarking vision-language models

Role
eval
Language
Python
Licence
Apache-2.0
Forks
768
Open issues
225
Last push
2026-09-17
Latest release
v0.1 · 2024-01-22
topicsevaluationmultimodalvqallmcomputer-visionpytorch