GH Repository · headroomlabs-ai
headroom
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
- stars
- 72,609
- 30-day movement
- +1,118373/day
- Related entries
- 60
- Connections
- 4
rustpython/uvmakedockerPythontokensprompt-engineeringcursortoken-optimizationairagcompressionmcp-serverproxycontext-engineeringopenaiagentcontext-windowtypescriptanthropicfastapimcpclaude-codelangchain
Headroom compresses tool outputs, logs, files, and RAG chunks before they reach an LLM, aiming to cut token usage while preserving answer quality. It ships as a library, a proxy, and an MCP server, with Python and Rust tooling.
Reach for it when agent or RAG token costs are dominated by bulky tool output and context payloads.
Use it to
- Compress JSON tool outputs before sending to an LLM
- Trim logs and file contents fed to coding agents
- Reduce RAG chunk token overhead
- Run it as an MCP server alongside Claude Code or Cursor
- Proxy LLM traffic to apply compression transparently
For Developers building LLM agents and RAG pipelines
- Role
- mcp-server
- Language
- Python
- Licence
- Apache-2.0
- Forks
- 5,578
- Open issues
- 325
- Last push
- 2026-09-17
- Latest release
- v0.2.15 · 2026-01-20
topicscompressiontoken-optimizationmcpcontext-engineeringagentsrag