BigHugger
GH Repository · opendataloader-project

opendataloader-pdf

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

stars
29,304
30-day movement
+9130/day
Related entries
61
Connections
2
a11ypdf-uanodepdf-extractionpdf-accessibilityocrjsonpdfJavaaidocument-parsinghtmlmarkdownpdf-convertertablesocr-recognitionragpdf-parsertagged-pdfaccessibilitybounding-boxeaa

opendataloader-pdf is an open-source Java PDF parser that converts PDFs into AI-ready formats such as JSON, Markdown, and HTML, with table extraction, OCR, and bounding-box output listed among its topics. It also targets PDF accessibility, including tagged PDF and PDF/UA handling.

Use it when you need to turn PDFs into structured data for RAG pipelines or automate PDF accessibility work.

Use it to

  • Parse PDFs into JSON or Markdown for RAG ingestion
  • Extract tables from PDF documents
  • Run OCR on scanned PDFs
  • Automate PDF accessibility tagging
  • Get bounding-box coordinates for extracted content

For Developers building document pipelines or accessibility tooling

Role
rag
Language
Java
Licence
Apache-2.0
Forks
2,791
Open issues
66
Last push
2026-09-17
Latest release
v0.0.1 · 2025-08-31
Skills shipped
1
topicspdfdocument-parsingragaccessibilityocrtables