nemotron speech streaming en 0.6b: transcribe.cpp GGUF GGUF conversions of nvidia/nemotron speech streaming en 0.6b for use with transcribe.cpp. Ported from upstream commit ef3bf40, pinned 2026 05 11. Validated against the NeMo reference at transcribe.cpp commit 12f1076 on 2026 05 11. Offline English speech to text with punctuation and capitalization. A 0.6B parameter cache aware streaming FastConformer encoder with an RNN T transducer decoder. Runs in offline mode. The encoder preserves the upstream att context size=[70, 13] (1.12s) cache aware attention mask end to end. transcribe.cpp does not yet expose a streaming session API. Downloads Quantization Download Size WER (LibriSpeech test clean, offline) : : F32 nemotron speech streaming en 0.6b F32.gguf 2.30 GB 2.31% F16 nemotron speech streaming en 0.6b F16.gguf 1.16 GB 2.31% Q8 0 nemotron speech streaming en 0.6b Q8 0.gguf 696 MB 2.31% Q6 K nemotron speech streaming en 0.6b Q6 K.gguf 573 MB 2.29% Q5 K M nemotron speech streaming en 0.6b Q5 K M.gguf 514 MB 2.34% Q4 K M nemotron speech streaming en 0.6b Q4 K M.gguf 453 MB 2.38% WER measured on the full LibriSpeech test clean split (2620 utterances) with greedy RNN T decoding. F32…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy