BigHugger
Deployability · 202,261 open models

Which models can you actually run?

Every leaderboard ranks models you call over an API. This one ranks models you can put on your own machine — by the runtime that loads them, the format they ship in, whether they're quantised, and whether the licence lets you sell what you build.

Read from the index 2026-09-18

Two different claims sit behind every bar, and they are kept apart everywhere on this page. Declared means a model card named the runtime. By format means the weights are in a container that runtime reads — which says the file will open, not that the architecture is implemented. A model is counted once if either is true. Nothing declares candle, burn or ort; no one writes a Rust runtime on a model card. That gap is the whole reason this page exists.

Reach, by runtime

🦀 burn113,832
mlx11,998
llama.cpp76,048
vllm110,181
091,046182,092 models

declared on the model cardnot declared, but ships a format it reads

Every bar but one is almost entirely light, which is the finding: for most runtimes the evidence is the file, not the card. MLX is the exception — 11,788 cards name it against 471 models shipping an npz, because mlx-community publishes converted weights under a name that says MLX rather than in the format that proves it. It is the one runtime where the card is the better evidence.

The most-used models mlx can load

Either signal
11,998
Declared
11,664
By format
486
Commercial use ok
8,169
Quantised build
9,315
ModelDownloadsParamsQuantLicenceTerms
pyannote/speaker-diarization-community-1 declared
automatic-speech-recognition
5,091,365CC-BY-4.0commercial ok
mlx-community/parakeet-tdt-0.6b-v3 declared
automatic-speech-recognition
1,828,956627MCC-BY-4.0commercial ok
OBLITERATUS/Qwen3.8-27B-OBLITERATED declared
text-generation
1,295,69728B2–16-bit
8 builds
Apache-2.0commercial ok
mlx-community/parakeet-tdt-0.6b-v2 declared
automatic-speech-recognition
1,279,723618MCC-BY-4.0commercial ok
prism-ml/Bonsai-27B-mlx-1bit declared
text-generation
1,057,2831.7BApache-2.0commercial ok
prism-ml/Ternary-Bonsai-27B-mlx-2bit declared
text-generation
1,047,63127B2-bitApache-2.0commercial ok
pnnbao-ump/VieNeu-TTS-v3-Turbo
text-to-speech
584,396131MApache-2.0commercial ok
lmstudio-community/DeepSeek-R1-0528-Qwen3-8B-MLX-4bit declared
text-generation
292,6128.2B4-bitMITcommercial ok
mlx-community/gpt-oss-20b-MXFP4-Q8 declared
text-generation
284,93321B4-bitApache-2.0commercial ok
lmstudio-community/DeepSeek-R1-0528-Qwen3-8B-MLX-8bit declared
text-generation
270,7768.2B8-bitMITcommercial ok
mlx-community/Qwen2.5-Coder-7B-Instruct-4bit declared
text-generation
258,7477.6B4-bitApache-2.0commercial ok
mlx-community/Qwen3-ASR-0.6B-8bit declared223,023782M8-bitApache-2.0commercial ok
orcarouter/Qwen3.8-27B-Uncensored-MLX declared
image-text-to-text
176,35927B4-bitApache-2.0commercial ok
pyannote-community/speaker-diarization-community-1 declared
automatic-speech-recognition
169,148CC-BY-4.0commercial ok
lmstudio-community/Qwen3-VL-8B-Instruct-MLX-5bit declared
image-text-to-text
148,2658.8B
base model
5-bitApache-2.0commercial ok
ampixa/sanoTTS
text-to-speech
141,331294,279
gguf size
GPL-3.0-onlycommercial ok
lmstudio-community/Qwen2.5-Coder-14B-Instruct-MLX-4bit declared
text-generation
128,09215B4-bitApache-2.0commercial ok
mlx-community/Qwen3.8-27B-4bit declared
image-text-to-text
126,32427B4-bitApache-2.0commercial ok
lmstudio-community/Qwen3-Coder-Next-MLX-8bit declared126,13180B8-bitApache-2.0commercial ok
lmstudio-community/Qwen3-Coder-Next-MLX-4bit declared115,17380B4-bitApache-2.0commercial ok
lmstudio-community/Qwen3-Coder-Next-MLX-6bit declared112,42080B6-bitApache-2.0commercial ok
lmstudio-community/Qwen2.5-Coder-14B-Instruct-MLX-8bit declared
text-generation
110,84215B8-bitApache-2.0commercial ok
lmstudio-community/Qwen3-VL-4B-Instruct-MLX-4bit declared
image-text-to-text
109,8144.4B4-bitApache-2.0commercial ok
mlx-community/Qwen3-0.6B-8bit declared
text-generation
104,372596M8-bitApache-2.0commercial ok
lmstudio-community/Qwen3-VL-4B-Instruct-MLX-6bit declared
image-text-to-text
103,1814.4B6-bitApache-2.0commercial ok
lmstudio-community/Qwen3-VL-4B-Instruct-MLX-8bit declared
image-text-to-text
103,0084.4B8-bitApache-2.0commercial ok
lmstudio-community/Qwen3-VL-4B-Instruct-MLX-5bit declared
image-text-to-text
102,4724.4B5-bitApache-2.0commercial ok
mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit declared
image-text-to-text
98,93124B4-bitApache-2.0commercial ok
lmstudio-community/Qwen3-VL-8B-Instruct-MLX-4bit declared
image-text-to-text
97,3728.8B
base model
4-bitApache-2.0commercial ok
lmstudio-community/Qwen3-VL-8B-Instruct-MLX-8bit declared
image-text-to-text
94,2378.8B
base model
8-bitApache-2.0commercial ok
lmstudio-community/Qwen3-VL-8B-Instruct-MLX-6bit declared
image-text-to-text
92,0228.8B
base model
6-bitApache-2.0commercial ok
mlx-community/all-MiniLM-L6-v2-4bit declared
sentence-similarity
91,99923M4-bitApache-2.0commercial ok
mlx-community/Qwen3-30B-A3B-Instruct-2507-4bit declared
text-generation
90,93731B4-bitApache-2.0commercial ok
mlx-community/Kokoro-82M-bf16 declared
text-to-speech
86,45082M
name
Apache-2.0commercial ok
mlx-community/Qwen3-8B-4bit declared
text-generation
81,6548.2B4-bitApache-2.0commercial ok
lmstudio-community/Qwen3-14B-MLX-4bit declared
text-generation
60,34015B4-bitApache-2.0commercial ok
Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed declared
text-generation
58,03327B4-bitApache-2.0commercial ok
lmstudio-community/Qwen3-14B-MLX-8bit declared
text-generation
55,23415B8-bitApache-2.0commercial ok
mlx-community/GLM-OCR-4bit declared
image-to-text
48,6991.1B4-bitMITcommercial ok
mlx-community/Qwen3.6-27B-4bit declared
image-text-to-text
48,42627B4-bitApache-2.0commercial ok
mlx-community/Qwen3-0.6B-4bit declared
text-generation
47,128596M4-bitApache-2.0commercial ok
lmstudio-community/Qwen2.5-Coder-32B-Instruct-MLX-4bit declared
text-generation
46,53033B4-bitApache-2.0commercial ok
lmstudio-community/Phi-4-mini-reasoning-MLX-4bit declared
text-generation
45,5733.8B4-bitMITcommercial ok
mlx-community/whisper-large-v3-mlx declared
automatic-speech-recognition
44,355MITcommercial ok
lmstudio-community/Qwen3-8B-MLX-4bit declared
text-generation
44,0038.2B4-bitApache-2.0commercial ok
mlx-community/Qwen3-4B-Instruct-2507-4bit declared
text-generation
43,9664.0B4-bitApache-2.0commercial ok
mlx-community/gemma-4-31b-it-4bit declared
image-text-to-text
43,60831B4-bitApache-2.0commercial ok
lmstudio-community/Qwen2.5-Coder-32B-Instruct-MLX-8bit declared
text-generation
43,59133B8-bitApache-2.0commercial ok
mlx-community/Qwen3.8-27B-8bit declared
image-text-to-text
41,72827B8-bitApache-2.0commercial ok
lmstudio-community/Qwen3-8B-MLX-8bit declared
text-generation
41,3618.2B8-bitApache-2.0commercial ok
mlx-community/GLM-5.2-mxfp4 declared
text-generation
40,317743B4-bitMITcommercial ok
KittenML/kitten-tts-nano-0.8-int840,1568-bitApache-2.0commercial ok
KittenML/kitten-tts-mini-0.838,287Apache-2.0commercial ok
mlx-community/Qwen3.8-27B-MTP-4bit declared
text-generation
36,462425M4-bitApache-2.0commercial ok
mlx-community/Mistral-7B-Instruct-v0.3-4bit declared33,6267.2B4-bitApache-2.0commercial ok
aufklarer/Silero-VAD-v5-MLX declared
voice-activity-detection
31,958309,121MITcommercial ok
mlx-community/Qwen3.6-35B-A3B-4bit declared
image-text-to-text
31,60335B4-bitApache-2.0commercial ok
lmstudio-community/Qwen3-1.7B-MLX-8bit declared
text-generation
30,4781.7B8-bitApache-2.0commercial ok
lmstudio-community/Qwen3-1.7B-MLX-4bit declared
text-generation
29,8891.7B4-bitApache-2.0commercial ok
lmstudio-community/Qwen3-VL-30B-A3B-Instruct-MLX-8bit declared
image-text-to-text
29,68831B
base model
8-bitApache-2.0commercial ok

Ordered by downloads, which on these 202,261 models is a real signal — 96% have a non-zero count. Parameter counts are resolved, not just read: the card first, then the base model's card, then the number in the name, then the GGUF file size, which is a per-parameter figure. That sizes 11,612 of the 11,998 models mlx can load; the source is printed under every count that did not come from the card itself. Quantisation is read the same way: the card's quant_bits or method where stated, otherwise the GGUF builds the repository actually ships, which is where a range like 2–16-bit and a build count come from.

The licence column resolved in full is on what you're allowed to ship. Everything here is queryable through the API.