BigHugger
GH Repository · Zipstack

unstract

LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows

stars
7,241
30-day movement
+83/day
Related entries
63
Connections
2
python/uvagent-appai-agentsdata-engineeringdocument-aigenerative-aiidppythonllmjson-extractionocrPythonmcp-serverpdf-extractionprompt-engineeringstructured-output

Unstract is an LLM-driven tool for extracting structured data from unstructured documents like PDFs, built to run as an API and inside ETL pipelines. It is written in Python, uses uv as the toolchain, ships with Claude directory configuration, and includes an MCP server component.

Reach for it when you need LLM-powered extraction of JSON-structured output from documents deployed behind APIs or embedded in data engineering workflows.

Use it to

  • Extract structured JSON from PDFs with LLMs
  • Run document extraction behind API endpoints
  • Integrate unstructured extraction into ETL pipelines
  • Expose extraction via its MCP server to agent clients
  • Apply OCR and prompt engineering to document processing

For Data engineers and developers building LLM document extraction pipelines

Role
agent-app
Language
Python
Licence
AGPL-3.0
Forks
716
Open issues
39
Last push
2026-09-17
Latest release
v0.1.0 · 2024-02-26
Skills shipped
3
topicsdocument-aillmunstructured-dataetlocrstructured-output