GH Repository · modelscope
evalscope
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
- stars
- 3,422
- 30-day movement
- +165/day
- Related entries
- 61
- Connections
- 2
pythonmakePythonvlmragperformancellmevalevaluation
EvalScope is a Python framework from ModelScope for evaluating and benchmarking large models, covering LLMs, VLMs, and AIGC models. It is customizable and also covers RAG evaluation per its topics.
Use it when you need a streamlined, configurable framework to benchmark large models rather than building evaluation pipelines yourself.
Use it to
- Benchmark LLM performance
- Evaluate vision-language models
- Assess RAG pipelines
- Customize evaluation workflows
- Run AIGC model evaluations
For Teams benchmarking and evaluating large models
- Role
- eval
- Language
- Python
- Licence
- Apache-2.0
- Forks
- 486
- Open issues
- 30
- Last push
- 2026-09-15
- Latest release
- v0.2.5 · 2024-04-02
- Skills shipped
- 1
topicsevaluationllmvlmragbenchmarkingpython