nemotron 3.5 asr streaming 0.6b: transcribe.cpp GGUF GGUF conversions of nvidia/nemotron 3.5 asr streaming 0.6b for use with transcribe.cpp. Ported from upstream commit 24b151a, pinned 2026 06 08. Validated against the NeMo reference at transcribe.cpp commit 909e94e on 2026 06 08. Multilingual speech to text across 32 supported language locales (the model's tokenizer recognizes 40, but 8 are adaptation ready and need fine tuning) with punctuation and capitalization. A 0.6B parameter cache aware streaming FastConformer encoder with a prompt conditioned RNN T transducer decoder; the target language is selected per call ( language en US, fr FR, de DE, ...) and an auto mode emits a tag. Ships both the offline path (att context size=[56, 13], 1.12s, headline accuracy) and runtime selectable chunked streaming ( stream chunk ms 1120 stream att right {0,3,6,13}). Downloads Quantization Download Size WER (FLEURS test en (en US), offline att context size=[56, 13]) : : F32 nemotron 3.5 asr streaming 0.6b F32.gguf 2.38 GB 7.97% F16 nemotron 3.5 asr streaming 0.6b F16.gguf 1.19 GB 7.97% Q8 0 nemotron 3.5 asr streaming 0.6b Q8 0.gguf 716 MB 7.88% Q6 K nemotron 3.5 asr streaming 0.6b Q6 K.gguf 59…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy