BigHugger
GH Repository · NanoNets

docext

An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)

stars
2,085
30-day movement
+10/day
Related entries
60
Connections
1
pythondockerocr-benchmarktable-extractionllm-ocrdocument-analysisvlmsPythononprem-visionocr-onpremiseextractiononpremiseonprem-ocrllmsmachine-learningdocument-data-extractionocrdocument-information-extractiondocumentnlpragunstructured-dataonprem

docext is an on-premises toolkit for extracting structured information from unstructured documents without OCR, converting documents to markdown, and benchmarking extraction performance. It is built around vision-language models and LLMs, and is associated with the IDP leaderboard at idp-leaderboard.org.

You need document extraction and table/markdown conversion that runs on your own infrastructure without an OCR pipeline.

Use it to

  • Extract structured data from unstructured documents on-premises
  • Convert documents to markdown
  • Extract tables from documents
  • Benchmark OCR-free extraction models
  • Feed extracted document data into RAG pipelines

For Teams doing on-premises document intelligence with LLMs

Role
rag
Language
Python
Licence
Apache-2.0
Forks
156
Open issues
21
Last push
2026-03-17
Latest release
v0.1.2 · 2025-04-05
topicsdocument-extractionocr-freevlmsonpremisesmarkdown-conversionbenchmarking