BigHugger
GH Repository · cactus-compute

cactus

Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.

stars
6,016
30-day movement
+186/day
Related entries
60
Connections
1
C++mobile-inferencearmquantizwhisperioson-device-aillamacppllmaillm-inferenceedge-aillmsragtransformeredgeandroidframeworkmobilesmartphonespeech

Cactus is a C++ inference engine providing quantization, kernels and a runtime for running LLMs and speech models on constrained devices. It targets mobiles, wearables, smart home devices and robots, with iOS and Android support and a RAG role.

You need on-device LLM and speech inference on ARM-based edge hardware rather than cloud APIs.

Use it to

  • Run LLM inference on smartphones
  • Deploy speech models with Whisper on-device
  • Build RAG features on mobile apps
  • Power AI in wearables and robots

For Mobile and edge developers building on-device AI apps

Role
rag
Language
C++
Licence
custom (licence file present)
Forks
502
Open issues
41
Last push
2026-09-08
Latest release
v1.0.0 · 2025-10-31
topicsllm-inferenceedge-aimobilequantizationragwhisper