GH Repository · yobix-ai
extractous
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
- stars
- 1,779
- 30-day movement
- +41/day
- Related entries
- 60
- Connections
- 1
Rustrustpdf-parserocrextractionpdfllmtikaunstructuredetlmachine-learningunstructured-datadata-pipelinesdocxnatural-language-processingragnlpetl-pipelines
Extractous is a Rust library for fast, efficient extraction of unstructured data from documents, with bindings for multiple languages. Its topic tags indicate support for PDF, DOCX, OCR, and Tika-style extraction, aimed at RAG and ETL workflows.
You reach for it when you need fast document-to-text extraction from a memory-efficient Rust core callable from several languages.
Use it to
- Parse PDFs to text for RAG pipelines
- Extract text from DOCX files
- Run OCR on scanned documents
- Build ETL data extraction steps
- Feed LLMs with cleaned document text
For Developers building RAG, NLP, or ETL data pipelines
- Role
- rag
- Language
- Rust
- Licence
- Apache-2.0
- Forks
- 96
- Open issues
- 24
- Last push
- 2024-12-21
topicsextractionpdfocrrustragetl