BigHugger
GH Repository · xberg-io

xberg

Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.

stars
9,314
30-day movement
+207/day
Related entries
60
Connections
2
typescriptswiftpython/uvnode/pnpmmcp-serverRustcsharpbunelixirwasmrubyrustdocument-intelligencemakeffipdfiummetadata-extractiontable-extractionpythontext-extractionjavaragphptesseract

xberg is a document intelligence tool with a Rust core that extracts text, metadata, images, tables, and structured data from 106 formats (140 extensions), plus code intelligence for 371 languages. It exposes this via a CLI, REST API, and MCP server, with bindings in fifteen languages.

You need uniform extraction from many document formats inside a pipeline or an agentic workflow, with an MCP server ready for LLM tool use.

Use it to

  • Extract text and metadata from PDFs for RAG indexing
  • Pull tables and structured data from office documents
  • Expose document extraction to agents via the MCP server
  • Serve extraction over a REST API from other applications
  • Analyze code structure across many languages

For Developers building RAG pipelines and document-processing or agent integrations

Role
mcp-server
Language
Rust
Licence
MIT
Forks
584
Open issues
5
Last push
2026-09-17
Latest release
benchmark-run-25624934491 · 2026-05-10
topicsdocument-intelligencetext-extractiontable-extractionmcp-serverrustrag