BigHugger
GH Repository · maximhq

bifrost

Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.

stars
8,130
30-day movement
+11338/day
Related entries
74
Connections
3
nixmakemodel-routertoken-managementload-balancingai-gatewaygateway-servicesGollm-gatewayguardrailsmcp-clientmcp-serverllmllm-costmcp-gatewayllm-observabilitygatewaygenerative-aillmops

Bifrost is a Go-based AI gateway for routing requests to 1000+ LLMs, with an adaptive load balancer, cluster mode, guardrails, and low-latency overhead (<100 µs at 5k RPS). It also acts as an MCP client and MCP server, and includes tooling for token management, cost tracking, and observability.

You need a high-throughput, self-hosted gateway that routes, balances, and guards LLM traffic across many models.

Use it to

  • Route LLM traffic across multiple providers and models
  • Balance and scale inference with cluster mode
  • Apply guardrails to requests and responses
  • Track token usage, cost, and observability metrics
  • Expose or consume MCP tools via the gateway

For Platform and LLMOps teams managing LLM traffic at scale

Role
mcp-client
Language
Go
Licence
Apache-2.0
Forks
1,226
Open issues
468
Last push
2026-09-17
Latest release
plugins/otel/v1.1.14 · 2026-01-22
Skills shipped
14
topicsllm-gatewayload-balancingguardrailsmodel-routermcpobservability