GH Repository · cactus-compute
cactus
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
- stars
- 6,016
- 30-day movement
- +186/day
- Related entries
- 60
- Connections
- 1
C++mobile-inferencearmquantizwhisperioson-device-aillamacppllmaillm-inferenceedge-aillmsragtransformeredgeandroidframeworkmobilesmartphonespeech
Cactus is a C++ inference engine providing quantization, kernels and a runtime for running LLMs and speech models on constrained devices. It targets mobiles, wearables, smart home devices and robots, with iOS and Android support and a RAG role.
You need on-device LLM and speech inference on ARM-based edge hardware rather than cloud APIs.
Use it to
- Run LLM inference on smartphones
- Deploy speech models with Whisper on-device
- Build RAG features on mobile apps
- Power AI in wearables and robots
For Mobile and edge developers building on-device AI apps
- Role
- rag
- Language
- C++
- Licence
- custom (licence file present)
- Forks
- 502
- Open issues
- 41
- Last push
- 2026-09-08
- Latest release
- v1.0.0 · 2025-10-31
topicsllm-inferenceedge-aimobilequantizationragwhisper