BigHugger
GH Repository · yfedoseev

pdf_oxide

The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.

stars
1,033
30-day movement
+103/day
Related entries
60
Connections
2
python/uvphpRustragpyo3pdf-to-textpdf-to-markdownmarkdownpdfpdf-editorpdf-librarymakepdf-parserrustllmpythonfastdocument-processingtext-extractiondata-extractionpdf-generationimage-extraction

pdf_oxide is a PDF library implemented in Rust with Python bindings via PyO3, covering text and image extraction, markdown conversion, and PDF creation and editing. The author reports 0.8ms mean processing time and a 100% pass rate on a 3,830-PDF test corpus.

Reach for it when you need fast, dual-language PDF parsing and generation without leaving Python or Rust.

Use it to

  • Extract text from PDFs
  • Convert PDFs to markdown
  • Extract embedded images
  • Create and edit PDFs
  • Prepare documents for RAG pipelines

For Python and Rust developers processing PDFs, especially for RAG

Role
rag
Language
Rust
Licence
Apache-2.0
Forks
129
Open issues
302
Last push
2026-09-18
Latest release
v0.1.0 · 2025-11-06
topicspdftext-extractionmarkdownrustpythonrag