BigHugger
GH Repository · k2-fsa

sherpa-onnx

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages

stars
14,817
30-day movement
+6722/day
Related entries
60
Connections
1
pythonjava/mavenotherC++lazarusmfcasrwindowsrisc-vandroidonnxmacosspeech-to-textdotnetcppswiftaarch64object-pascaliostext-to-speechlinuxraspberry-piarm32vits

sherpa-onnx is a C++ speech toolkit built on next-gen Kaldi and onnxruntime that runs speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD fully offline. It targets a wide range of platforms including Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, and NPU devices, and exposes bindings in 12 programming languages plus websocket server/client support.

Use it when you need offline speech recognition or synthesis that runs on embedded and mobile hardware, not just cloud servers.

Use it to

  • Run on-device speech recognition on Android or iOS apps
  • Deploy text-to-speech on Raspberry Pi or RISC-V boards
  • Add speaker diarization to offline audio pipelines
  • Serve ASR/TTS over a local websocket server
  • Integrate speech features via C#, Java, Swift, or Python bindings

For Developers building offline speech apps across mobile, embedded, and server platforms

Role
other
Language
C++
Licence
Apache-2.0
Forks
1,701
Open issues
563
Last push
2026-09-17
Latest release
v1.0 · 2022-10-14
topicsspeech-to-texttext-to-speechonnxembeddedspeaker-diarizationvad