BigHugger
GH Repository · yobix-ai

extractous

Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.

stars
1,779
30-day movement
+41/day
Related entries
60
Connections
1
Rustrustpdf-parserocrextractionpdfllmtikaunstructuredetlmachine-learningunstructured-datadata-pipelinesdocxnatural-language-processingragnlpetl-pipelines

Extractous is a Rust library for fast, efficient extraction of unstructured data from documents, with bindings for multiple languages. Its topic tags indicate support for PDF, DOCX, OCR, and Tika-style extraction, aimed at RAG and ETL workflows.

You reach for it when you need fast document-to-text extraction from a memory-efficient Rust core callable from several languages.

Use it to

  • Parse PDFs to text for RAG pipelines
  • Extract text from DOCX files
  • Run OCR on scanned documents
  • Build ETL data extraction steps
  • Feed LLMs with cleaned document text

For Developers building RAG, NLP, or ETL data pipelines

Role
rag
Language
Rust
Licence
Apache-2.0
Forks
96
Open issues
24
Last push
2024-12-21
topicsextractionpdfocrrustragetl