parakeet unified en 0.6b: transcribe.cpp GGUF GGUF conversions of nvidia/parakeet unified en 0.6b for use with transcribe.cpp. Ported from upstream commit d4ac992, pinned 2026 05 10. Validated against the NeMo reference at transcribe.cpp commit 42528dd on 2026 05 10. Offline English speech to text with punctuation and capitalization. A 0.6B parameter FastConformer encoder with an RNN T transducer decoder, trained as a 'unified' streaming/offline model. This port runs the model in offline mode only — streaming attention contexts are present in the GGUF but transcribe.cpp does not yet expose a streaming entry. Downloads Quantization Download Size WER (LibriSpeech test clean, offline) : : F32 parakeet unified en 0.6b F32.gguf 2.47 GB 1.59% F16 parakeet unified en 0.6b F16.gguf 1.24 GB 1.59% Q8 0 parakeet unified en 0.6b Q8 0.gguf 731 MB 1.60% Q6 K parakeet unified en 0.6b Q6 K.gguf 602 MB 1.61% Q5 K M parakeet unified en 0.6b Q5 K M.gguf 541 MB 1.58% Q4 K M parakeet unified en 0.6b Q4 K M.gguf 477 MB 1.62% WER measured on the full LibriSpeech test clean split (2620 utterances) with greedy RNN T decoding and no external LM. F32 reference baseline: 1.59%. NVIDIA's self reported number o…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy