canary 180m flash: transcribe.cpp GGUF GGUF conversions of nvidia/canary 180m flash for use with transcribe.cpp. Ported from upstream commit b12ab41, pinned 2026 05 08. Validated against the NeMo reference at transcribe.cpp commit db53eda on 2026 05 08. Offline multilingual speech to text and translation. A 182M parameter multitask AED with a 17 layer FastConformer encoder and a 4 layer Transformer decoder. Supports automatic speech recognition in English, German, Spanish, and French, and bidirectional EN↔{DE, ES, FR} translation. Takes a 16 kHz mono WAV and produces a transcript. Not a streaming model; word/segment timestamps are upstream experimental and not exposed in the v1 port. Downloads Quantization Download Size WER (LibriSpeech test clean) : : F32 canary 180m flash F32.gguf 721 MB 1.94% F16 canary 180m flash F16.gguf 364 MB 1.94% Q8 0 canary 180m flash Q8 0.gguf 208 MB 1.93% Q6 K canary 180m flash Q6 K.gguf 168 MB 1.93% Q5 K M canary 180m flash Q5 K M.gguf 151 MB 1.90% Q4 K M canary 180m flash Q4 K M.gguf 133 MB 1.93% WER measured on the full LibriSpeech test clean split (2620 utterances) with greedy decoding and no external LM. F32 reference baseline: 1.94%. On the same w…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy