BigHugger
GH Repository · PaddlePaddle

PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

stars
89,698
30-day movement
+29197/day
Related entries
62
Connections
1
Pythonocrchineseocrpdf-parserpythonpdf2markdownpp-ocrdocument-parsingdocument-translationpdf-extractor-ragkiepp-structureai4scienceragpaddleocr-vl

PaddleOCR is a Python OCR toolkit built on PaddlePaddle that converts PDFs and images into structured data, supporting 100+ languages. Its topics indicate capabilities spanning OCR, document parsing, PDF-to-markdown conversion, key information extraction, and RAG-oriented document extraction.

Use it when you need to turn document images or PDFs into structured text that LLMs or RAG pipelines can consume.

Use it to

  • Extract text from scanned documents and images
  • Convert PDFs to markdown for RAG ingestion
  • Parse document structure from mixed-layout files
  • Perform key information extraction on documents
  • Recognize text across 100+ languages

For Developers building document-processing and RAG pipelines

Role
rag
Language
Python
Licence
Apache-2.0
Forks
11,349
Open issues
167
Last push
2026-09-16
Latest release
v1.1.0 · 2020-09-27
Skills shipped
2
topicsocrdocument-parsingpdf-parserragchineseocrkie