BigHugger
GH Repository · Unstructured-IO

unstructured

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

stars
15,439
30-day movement
starts with the next reading
Related entries
60
Connections
1
python/uvpython/condamakedockerpythonllmdocument-parserdonutdocument-image-analysisdeep-learningdocument-image-processingdocument-parsingHTMLmachine-learningdocxdata-pipelinesnlpnatural-language-processingocrinformation-retrievalagent-frameworkpdf-to-textpreprocessingml

Unstructured is an open-source ETL library that converts complex documents such as PDFs and DOCX files into clean, structured formats suitable for language models. The repository also points to a commercial Platform product offering partitioning, chunking, embedding and enrichments for production workflows.

You reach for it when you need to turn raw documents into LLM-ready structured data without writing your own parsing pipeline.

Use it to

  • Partition PDFs into structured elements
  • Convert DOCX documents to text or JSON
  • Preprocess documents for LLM ingestion
  • Chunk and prepare document data for RAG pipelines

For Developers building LLM and document-processing pipelines

Role
agent-framework
Language
HTML
Licence
Apache-2.0
Forks
1,329
Open issues
183
Last push
2026-09-15
Latest release
0.2.1 · 2022-10-21
topicsdocument-parsingetlpdf-to-textocrllmpreprocessing