BigHugger
GH Repository · beir-cellar

beir

A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.

stars
2,291
30-day movement
+41/day
Related entries
60
Connections
1
Pythonpythonllmpytorchzero-shot-retrievalbenchmarkdeep-learningcolbertsbertbertinformation-retrievalsentence-transformersquestion-generationragnlpretrievaldprevalpassage-retrievalelasticsearchretrieval-modelsdataset

BEIR is a heterogeneous benchmark for information retrieval that packages 15+ diverse IR datasets behind a common evaluation interface. It lets you score retrieval models (BERT, DPR, ColBERT, SBERT, BM25/Elasticsearch) on zero-shot retrieval tasks.

You need one consistent harness to compare retrieval models across many datasets instead of wiring up each benchmark separately.

Use it to

  • Evaluate a retrieval model on 15+ IR datasets with one pipeline
  • Run zero-shot retrieval comparisons across benchmarks
  • Benchmark dense retrievers like DPR, ColBERT, SBERT
  • Score sparse baselines such as BM25 via Elasticsearch
  • Build RAG retrieval evaluation into your workflow

For IR and NLP researchers, engineers evaluating retrieval or RAG systems

Role
eval
Language
Python
Licence
Apache-2.0
Forks
251
Open issues
66
Last push
2025-10-16
Latest release
v0.2.0 · 2021-07-06
topicsinformation-retrievalbenchmarkevaluationzero-shot-retrievalsentence-transformersrag