SenseVoiceSmall: transcribe.cpp GGUF GGUF conversions of FunAudioLLM/SenseVoiceSmall for use with transcribe.cpp. Ported from upstream commit 3eb3b4eeffc2f2dde6051b853983753db33e35c3, pinned 2026 05 06. Validated against the FunASR reference at transcribe.cpp commit f094d28 on 2026 05 06. Offline multilingual speech to text in Chinese, Cantonese, English, Japanese, and Korean. A 234M parameter SAN M encoder with a single CTC head over a 25,055 token SentencePiece vocabulary. Takes a 16 kHz mono WAV (capped at 30 seconds per call, per upstream's direct inference contract) and produces a transcript. Not a streaming model, no translation, no built in long form chunking. The same CTC head also emits language ID, simple emotion labels, audio event tags, and an inverse text normalization flag — opt in via raw tokens and itn . Downloads Quantization Download Size WER (LibriSpeech test clean) : : F32 SenseVoiceSmall F32.gguf 893 MB 3.13% F16 SenseVoiceSmall F16.gguf 449 MB 3.13% Q8 0 SenseVoiceSmall Q8 0.gguf 241 MB 3.13% Q6 K SenseVoiceSmall Q6 K.gguf 187 MB 3.14% Q5 K M SenseVoiceSmall Q5 K M.gguf 164 MB 3.18% Q4 K M SenseVoiceSmall Q4 K M.gguf 139 MB 3.45% WER measured on the full Libri…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy