BigHugger
GH Repository · enoch3712

ExtractThinker

ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.

stars
1,598
30-day movement
starts with the next reading
Related entries
60
Connections
1
python/poetryagent-frameworkpdf-to-textainlpocropenaillmpythonmachine-learningdocument-image-analysisdocument-intelligencedocument-parsingPythondocument-processingpdflangchain

ExtractThinker is a Python Document Intelligence library for LLMs that provides ORM-style interaction for building document workflows. Its topics indicate support for OCR, PDF-to-text parsing, and document processing, with references to OpenAI and LangChain.

Reach for it when you want structured, ORM-like handling of document extraction pipelines powered by LLMs.

Use it to

  • Parse PDFs into structured data with LLMs
  • Run OCR on document images
  • Build document processing workflows
  • Integrate extraction with LangChain pipelines

For Python developers building LLM-based document extraction pipelines

Role
agent-framework
Language
Python
Licence
Apache-2.0
Forks
153
Last push
2026-09-16
Latest release
v0.0.2 · 2024-05-20
topicsdocument-intelligencellmocrpdfpythonlangchain