sherpa-onnx
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
- stars
- 14,817
- 30-day movement
- +6722/day
- Related entries
- 60
- Connections
- 1
sherpa-onnx is a C++ speech toolkit built on next-gen Kaldi and onnxruntime that runs speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD fully offline. It targets a wide range of platforms including Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, and NPU devices, and exposes bindings in 12 programming languages plus websocket server/client support.
Use it when you need offline speech recognition or synthesis that runs on embedded and mobile hardware, not just cloud servers.
Use it to
- Run on-device speech recognition on Android or iOS apps
- Deploy text-to-speech on Raspberry Pi or RISC-V boards
- Add speaker diarization to offline audio pipelines
- Serve ASR/TTS over a local websocket server
- Integrate speech features via C#, Java, Swift, or Python bindings
For Developers building offline speech apps across mobile, embedded, and server platforms
- Role
- other
- Language
- C++
- Licence
- Apache-2.0
- Forks
- 1,701
- Open issues
- 563
- Last push
- 2026-09-17
- Latest release
- v1.0 · 2022-10-14