BigHugger
GH Repository · modelscope

evalscope

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

stars
3,422
30-day movement
+165/day
Related entries
61
Connections
2
pythonmakePythonvlmragperformancellmevalevaluation

EvalScope is a Python framework from ModelScope for evaluating and benchmarking large models, covering LLMs, VLMs, and AIGC models. It is customizable and also covers RAG evaluation per its topics.

Use it when you need a streamlined, configurable framework to benchmark large models rather than building evaluation pipelines yourself.

Use it to

  • Benchmark LLM performance
  • Evaluate vision-language models
  • Assess RAG pipelines
  • Customize evaluation workflows
  • Run AIGC model evaluations

For Teams benchmarking and evaluating large models

Role
eval
Language
Python
Licence
Apache-2.0
Forks
486
Open issues
30
Last push
2026-09-15
Latest release
v0.2.5 · 2024-04-02
Skills shipped
1
topicsevaluationllmvlmragbenchmarkingpython