GH Repository · maximhq
bifrost
Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.
- stars
- 8,130
- 30-day movement
- +11338/day
- Related entries
- 74
- Connections
- 3
nixmakemodel-routertoken-managementload-balancingai-gatewaygateway-servicesGollm-gatewayguardrailsmcp-clientmcp-serverllmllm-costmcp-gatewayllm-observabilitygatewaygenerative-aillmops
Bifrost is a Go-based AI gateway for routing requests to 1000+ LLMs, with an adaptive load balancer, cluster mode, guardrails, and low-latency overhead (<100 µs at 5k RPS). It also acts as an MCP client and MCP server, and includes tooling for token management, cost tracking, and observability.
You need a high-throughput, self-hosted gateway that routes, balances, and guards LLM traffic across many models.
Use it to
- Route LLM traffic across multiple providers and models
- Balance and scale inference with cluster mode
- Apply guardrails to requests and responses
- Track token usage, cost, and observability metrics
- Expose or consume MCP tools via the gateway
For Platform and LLMOps teams managing LLM traffic at scale
- Role
- mcp-client
- Language
- Go
- Licence
- Apache-2.0
- Forks
- 1,226
- Open issues
- 468
- Last push
- 2026-09-17
- Latest release
- plugins/otel/v1.1.14 · 2026-01-22
- Skills shipped
- 14
topicsllm-gatewayload-balancingguardrailsmodel-routermcpobservability