GH Repository · NanoNets
docext
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
- stars
- 2,085
- 30-day movement
- +10/day
- Related entries
- 60
- Connections
- 1
pythondockerocr-benchmarktable-extractionllm-ocrdocument-analysisvlmsPythononprem-visionocr-onpremiseextractiononpremiseonprem-ocrllmsmachine-learningdocument-data-extractionocrdocument-information-extractiondocumentnlpragunstructured-dataonprem
docext is an on-premises toolkit for extracting structured information from unstructured documents without OCR, converting documents to markdown, and benchmarking extraction performance. It is built around vision-language models and LLMs, and is associated with the IDP leaderboard at idp-leaderboard.org.
You need document extraction and table/markdown conversion that runs on your own infrastructure without an OCR pipeline.
Use it to
- Extract structured data from unstructured documents on-premises
- Convert documents to markdown
- Extract tables from documents
- Benchmark OCR-free extraction models
- Feed extracted document data into RAG pipelines
For Teams doing on-premises document intelligence with LLMs
- Role
- rag
- Language
- Python
- Licence
- Apache-2.0
- Forks
- 156
- Open issues
- 21
- Last push
- 2026-03-17
- Latest release
- v0.1.2 · 2025-04-05
topicsdocument-extractionocr-freevlmsonpremisesmarkdown-conversionbenchmarking