BigHugger
GH Repository · waybarrios

vllm-mlx

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

stars
1,571
30-day movement
+72/day
Related entries
60
Connections
1
Pythonmcpopenaiinference-serverapple-siliconpythonllmtool-callingmlxopenai-compatibleopenai-apimacosvision-language-modelclaude-codevllmmultimodal-aispeech-to-textanthropicmcp-serveranthropic-apilocal-llmtext-to-speechcontinuous-batching
Role
mcp-server
Language
Python
Licence
Apache-2.0
Forks
220
Open issues
67
Last push
2026-09-06
Latest release
v0.1.0 · 2026-01-06