GH Repository · beir-cellar
beir
A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.
- stars
- 2,291
- 30-day movement
- +41/day
- Related entries
- 60
- Connections
- 1
Pythonpythonllmpytorchzero-shot-retrievalbenchmarkdeep-learningcolbertsbertbertinformation-retrievalsentence-transformersquestion-generationragnlpretrievaldprevalpassage-retrievalelasticsearchretrieval-modelsdataset
BEIR is a heterogeneous benchmark for information retrieval that packages 15+ diverse IR datasets behind a common evaluation interface. It lets you score retrieval models (BERT, DPR, ColBERT, SBERT, BM25/Elasticsearch) on zero-shot retrieval tasks.
You need one consistent harness to compare retrieval models across many datasets instead of wiring up each benchmark separately.
Use it to
- Evaluate a retrieval model on 15+ IR datasets with one pipeline
- Run zero-shot retrieval comparisons across benchmarks
- Benchmark dense retrievers like DPR, ColBERT, SBERT
- Score sparse baselines such as BM25 via Elasticsearch
- Build RAG retrieval evaluation into your workflow
For IR and NLP researchers, engineers evaluating retrieval or RAG systems
- Role
- eval
- Language
- Python
- Licence
- Apache-2.0
- Forks
- 251
- Open issues
- 66
- Last push
- 2025-10-16
- Latest release
- v0.2.0 · 2021-07-06
topicsinformation-retrievalbenchmarkevaluationzero-shot-retrievalsentence-transformersrag