Which models can you actually run?
Every leaderboard ranks models you call over an API. This one ranks models you can put on your own machine — by the runtime that loads them, the format they ship in, whether they're quantised, and whether the licence lets you sell what you build.
Read from the index 2026-09-18Two different claims sit behind every bar, and they are kept apart everywhere on this page. Declared means a model card named the runtime. By format means the weights are in a container that runtime reads — which says the file will open, not that the architecture is implemented. A model is counted once if either is true. Nothing declares candle, burn or ort; no one writes a Rust runtime on a model card. That gap is the whole reason this page exists.
Reach, by runtime
declared on the model cardnot declared, but ships a format it reads
Every bar but one is almost entirely light, which is the finding: for most runtimes the evidence is the file, not the card. MLX is the exception — 11,788 cards name it against 471 models shipping an npz, because mlx-community publishes converted weights under a name that says MLX rather than in the format that proves it. It is the one runtime where the card is the better evidence.
The most-used models llama.cpp can load
- Either signal
- 76,048
- Declared
- 973
- By format
- 76,013
- Commercial use ok
- 51,565
- Quantised build
- 72,603
| Model | Downloads | Params | Quant | Licence | Terms |
|---|---|---|---|---|---|
| unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF text-generation | 12,866,779 | 31B base model | 1–16-bit 26 builds | Apache-2.0 | commercial ok |
| unsloth/Qwen3.8-27B-GGUF | 8,856,150 | 28B base model | 1–16-bit 26 builds | Apache-2.0 | commercial ok |
| ornith-ai/Ornith-1.5-9B-GGUF text-generation | 5,320,513 | 9.0B name | 4–16-bit 5 builds | MIT | commercial ok |
| ornith-ai/Ornith-1.5-35B-A3B-GGUF text-generation | 4,378,772 | 35B name | 4–16-bit 5 builds | MIT | commercial ok |
| audio-cpp/audio.cpp-gguf text-to-speech | 3,800,350 | 1.5B gguf size | 4–32-bit 8 builds | MIT | commercial ok |
| ornith-ai/Ornith-1.0-9B-GGUF text-generation | 3,514,831 | 9.0B name | 4–16-bit 5 builds | MIT | commercial ok |
| cdiamond/Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF image-text-to-text | 3,307,312 | 28B base model | — | Apache-2.0 | commercial ok |
| huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF image-text-to-text | 2,828,411 | 28B base model | 2–16-bit 26 builds | Apache-2.0 | commercial ok |
| JonathanColetti/Qwen3.8-27B-Uncensored-GGUF declared text-generation | 2,688,909 | 28B base model | 2–16-bit 8 builds | Apache-2.0 | commercial ok |
| mixedbread-ai/mxbai-embed-large-v1 feature-extraction | 2,523,873 | 335M | — | Apache-2.0 | commercial ok |
| HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF image-text-to-text | 2,387,042 | 28B base model | 2–16-bit 11 builds | Apache-2.0 | commercial ok |
| lmstudio-community/Qwen3.8-27B-GGUF | 2,328,522 | 28B base model | 4–16-bit 4 builds | Apache-2.0 | commercial ok |
| ornith-ai/Ornith-1.0-35B-GGUF text-generation | 2,265,459 | 35B name | 4–16-bit 5 builds | MIT | commercial ok |
| mudler/Laguna-XS-2.1-APEX-GGUF | 2,203,383 | 33B base model | — | — | unknown |
| HauhauCS/Gemma-4-E4B-Uncensored-HauhauCS-Aggressive image-text-to-text | 2,133,271 | 7.5B gguf size | 2–16-bit 12 builds | — | conditional |
| antirez/deepseek-v4-gguf text-generation | 2,031,158 | 291B base model | 4-bit 3 builds | MIT | commercial ok |
| 0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF text-generation | 1,972,780 | 28B base model | 1–16-bit 25 builds | Apache-2.0 | commercial ok |
| handy-computer/nemotron-3.5-asr-streaming-0.6b-gguf automatic-speech-recognition | 1,830,030 | 638M base model | 4–32-bit 6 builds | — | unknown |
| unsloth/Qwen3.5-9B-GGUF image-text-to-text | 1,596,048 | 9.7B base model | 2–32-bit 24 builds | Apache-2.0 | commercial ok |
| handy-computer/parakeet-unified-en-0.6b-gguf automatic-speech-recognition | 1,562,280 | 600M name | 4–32-bit 6 builds | CC-BY-4.0 | commercial ok |
| unsloth/Inkling-Small-GGUF image-text-to-text | 1,435,922 | 266B base model | 1–32-bit 24 builds | Apache-2.0 | commercial ok |
| ggml-org/gemma-4-E4B-it-GGUF any-to-any | 1,430,536 | 8.0B base model | 4–16-bit 3 builds | Apache-2.0 | commercial ok |
| unsloth/Qwen3.8-Flash-Next-GGUF image-text-to-text | 1,371,824 | 180B base model | 1–16-bit 13 builds | — | unknown |
| mudler/KAT-Coder-V2.5-Dev-APEX-GGUF | 1,364,645 | 35B base model | — | Apache-2.0 | commercial ok |
| ggml-org/Qwen3.8-27B-GGUF image-text-to-text | 1,354,640 | 28B base model | 4–16-bit 4 builds | Apache-2.0 | commercial ok |
| DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF image-text-to-text | 1,331,900 | 9.7B base model | 2–32-bit 13 builds | Apache-2.0 | commercial ok |
| unsloth/gemma-4-12B-it-qat-GGUF any-to-any | 1,328,589 | 12B base model | 4–32-bit 6 builds | Apache-2.0 | commercial ok |
| ornith-ai/Ornith-1.5-397B-GGUF text-generation | 1,327,991 | 397B name | 4–16-bit 5 builds | MIT | commercial ok |
| OBLITERATUS/Qwen3.8-27B-OBLITERATED text-generation | 1,295,697 | 28B | 2–16-bit 8 builds | Apache-2.0 | commercial ok |
| unsloth/Qwen3.6-35B-A3B-GGUF image-text-to-text | 1,268,083 | 36B base model | 1–32-bit 25 builds | Apache-2.0 | commercial ok |
| unsloth/Qwen3.6-27B-MTP-GGUF image-text-to-text | 1,230,293 | 28B base model | 2–32-bit 24 builds | Apache-2.0 | commercial ok |
| LiquidAI/LFM2.5-2.6B-GGUF text-generation | 1,182,734 | 2.7B base model | 4–16-bit 7 builds | — | unknown |
| unsloth/Qwen3.6-27B-GGUF image-text-to-text | 1,135,646 | 28B base model | 2–32-bit 24 builds | Apache-2.0 | commercial ok |
| datalab-to/surya-ocr-2-gguf image-text-to-text | 1,127,344 | 632M gguf size | — | — | conditional |
| mudler/ced-gguf audio-classification | 1,074,342 | 86M base model | 8–32-bit 3 builds | Apache-2.0 | commercial ok |
| HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive image-text-to-text | 1,068,618 | 36B base model | 2–16-bit 12 builds | Apache-2.0 | commercial ok |
| Abiray/MiniMax-H3-GGUF image-to-video | 1,055,188 | 33B base model | 3–8-bit 10 builds | — | unknown |
| DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF image-text-to-text | 1,049,586 | 28B base model | 2–32-bit 12 builds | Apache-2.0 | commercial ok |
| legraphista/glm-4-9b-chat-IMat-GGUF text-generation | 1,011,242 | 9.4B base model | 1–16-bit 24 builds | — | unknown |
| rippertnt/HyperCLOVAX-SEED-Text-Instruct-1.5B-Q4_K_M-GGUF | 992,905 | 1.6B base model | 4-bit | — | unknown |
| bartowski/endless-frontier_BigBang-v1-GGUF image-text-to-text | 986,052 | 36B base model | 2–16-bit 28 builds | Apache-2.0 | commercial ok |
| ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF image-text-to-text | 956,964 | 28B base model | 2–16-bit 5 builds | Apache-2.0 | commercial ok |
| handy-computer/cohere-transcribe-03-2026-gguf automatic-speech-recognition | 955,900 | 2.1B base model | 4–16-bit 6 builds | Apache-2.0 | commercial ok |
| LocalAI-io/privacy-filter-nemotron-GGUF token-classification | 953,658 | 1.4B base model | 8–16-bit 2 builds | Apache-2.0 | commercial ok |
| unsloth/Qwen3.6-35B-A3B-MTP-GGUF image-text-to-text | 948,190 | 36B base model | 1–32-bit 23 builds | Apache-2.0 | commercial ok |
| unsloth/Qwen3.5-4B-GGUF image-text-to-text | 923,154 | 4.7B base model | 2–32-bit 24 builds | Apache-2.0 | commercial ok |
| DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF image-text-to-text | 907,640 | 28B base model | 2–32-bit 13 builds | Apache-2.0 | commercial ok |
| bartowski/XYZAILab_XYZ-Aquila-mini-GGUF image-text-to-text | 883,330 | 36B base model | 2–16-bit 28 builds | Apache-2.0 | commercial ok |
| HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive | 867,389 | 9.7B base model | 4–16-bit 4 builds | Apache-2.0 | commercial ok |
| nvidia/parakeet-ctc-1.1b automatic-speech-recognition | 841,226 | 1.1B | 8-bit | CC-BY-4.0 | commercial ok |
| unsloth/MiniMax-H3-GGUF image-text-to-video | 837,860 | 33B base model | 2–8-bit 10 builds | — | unknown |
| FINAL-Bench/POCKET-35B-GGUF declared text-generation | 825,835 | 35B base model | 1–4-bit 4 builds | Apache-2.0 | commercial ok |
| google/gemma-4-12B-it-qat-q4_0-gguf any-to-any | 793,355 | 12B base model | 4-bit | Apache-2.0 | commercial ok |
| empero-ai/Qwen3.8-4B-Distill-GGUF text-generation | 771,434 | 4.7B base model | 4–16-bit 5 builds | Apache-2.0 | commercial ok |
| empero-ai/Qwen3.8-2B-Distill-GGUF text-generation | 749,203 | 2.3B base model | 4–16-bit 5 builds | Apache-2.0 | commercial ok |
| nvidia/nemotron-3.5-asr-streaming-0.6b automatic-speech-recognition | 742,455 | 638M | 8-bit | — | unknown |
| QuantStack/Wan2.2-T2V-A14B-GGUF text-to-video | 724,592 | 14B gguf size | 2–8-bit 13 builds | Apache-2.0 | commercial ok |
| empero-ai/Qwen3.8-9B-Distill-GGUF text-generation | 699,445 | 9.7B base model | 4–16-bit 5 builds | Apache-2.0 | commercial ok |
| google/gemma-4-E4B-it-qat-q4_0-gguf any-to-any | 693,070 | 8.0B base model | 4-bit | Apache-2.0 | commercial ok |
| yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF text-generation | 689,635 | 12B base model | 3–16-bit 6 builds | Apache-2.0 | commercial ok |
Ordered by downloads, which on these 202,261 models is a real signal — 96% have a non-zero count. Parameter counts are resolved, not just read: the card first, then the base model's card, then the number in the name, then the GGUF file size, which is a per-parameter figure. That sizes 76,023 of the 76,048 models llama.cpp can load; the source is printed under every count that did not come from the card itself. Quantisation is read the same way: the card's quant_bits or method where stated, otherwise the GGUF builds the repository actually ships, which is where a range like 2–16-bit and a build count come from.
The licence column resolved in full is on what you're allowed to ship. Everything here is queryable through the API.