GH Repository · yfedoseev
pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
- stars
- 1,033
- 30-day movement
- +103/day
- Related entries
- 60
- Connections
- 2
python/uvphpRustragpyo3pdf-to-textpdf-to-markdownmarkdownpdfpdf-editorpdf-librarymakepdf-parserrustllmpythonfastdocument-processingtext-extractiondata-extractionpdf-generationimage-extraction
pdf_oxide is a PDF library implemented in Rust with Python bindings via PyO3, covering text and image extraction, markdown conversion, and PDF creation and editing. The author reports 0.8ms mean processing time and a 100% pass rate on a 3,830-PDF test corpus.
Reach for it when you need fast, dual-language PDF parsing and generation without leaving Python or Rust.
Use it to
- Extract text from PDFs
- Convert PDFs to markdown
- Extract embedded images
- Create and edit PDFs
- Prepare documents for RAG pipelines
For Python and Rust developers processing PDFs, especially for RAG
- Role
- rag
- Language
- Rust
- Licence
- Apache-2.0
- Forks
- 129
- Open issues
- 302
- Last push
- 2026-09-18
- Latest release
- v0.1.0 · 2025-11-06
topicspdftext-extractionmarkdownrustpythonrag