BigHugger
GH Repository · headroomlabs-ai

headroom

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.

stars
72,242
30-day movement
+1,118373/day
Related entries
60
Connections
4
rustpython/uvmakedockerPythontokensprompt-engineeringcursortoken-optimizationairagcompressionmcp-serverproxycontext-engineeringopenaiagentcontext-windowtypescriptanthropicfastapimcpclaude-codelangchain

Headroom compresses tool outputs, logs, files, and RAG chunks before they reach an LLM, aiming to cut token usage while preserving answer quality. It ships as a library, a proxy, and an MCP server, with Python and Rust tooling.

Reach for it when agent or RAG token costs are dominated by bulky tool output and context payloads.

Use it to

  • Compress JSON tool outputs before sending to an LLM
  • Trim logs and file contents fed to coding agents
  • Reduce RAG chunk token overhead
  • Run it as an MCP server alongside Claude Code or Cursor
  • Proxy LLM traffic to apply compression transparently

For Developers building LLM agents and RAG pipelines

Role
mcp-server
Language
Python
Licence
Apache-2.0
Forks
5,531
Open issues
318
Last push
2026-09-15
Latest release
v0.2.15 · 2026-01-20
topicscompressiontoken-optimizationmcpcontext-engineeringagentsrag