GH Repository · PaddlePaddle
PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
- stars
- 89,698
- 30-day movement
- +29197/day
- Related entries
- 62
- Connections
- 1
Pythonocrchineseocrpdf-parserpythonpdf2markdownpp-ocrdocument-parsingdocument-translationpdf-extractor-ragkiepp-structureai4scienceragpaddleocr-vl
PaddleOCR is a Python OCR toolkit built on PaddlePaddle that converts PDFs and images into structured data, supporting 100+ languages. Its topics indicate capabilities spanning OCR, document parsing, PDF-to-markdown conversion, key information extraction, and RAG-oriented document extraction.
Use it when you need to turn document images or PDFs into structured text that LLMs or RAG pipelines can consume.
Use it to
- Extract text from scanned documents and images
- Convert PDFs to markdown for RAG ingestion
- Parse document structure from mixed-layout files
- Perform key information extraction on documents
- Recognize text across 100+ languages
For Developers building document-processing and RAG pipelines
- Role
- rag
- Language
- Python
- Licence
- Apache-2.0
- Forks
- 11,349
- Open issues
- 167
- Last push
- 2026-09-16
- Latest release
- v1.1.0 · 2020-09-27
- Skills shipped
- 2
topicsocrdocument-parsingpdf-parserragchineseocrkie