canary 1b flash: transcribe.cpp GGUF GGUF conversions of nvidia/canary 1b flash for use with transcribe.cpp. Ported from upstream commit a9a55e0, pinned 2026 05 08. Validated against the NeMo reference at transcribe.cpp commit db53eda on 2026 05 08. Offline multilingual speech to text and translation. An 883M parameter multitask AED with a 32 layer FastConformer encoder and a 4 layer Transformer decoder. Supports automatic speech recognition in English, German, Spanish, and French, and bidirectional EN↔{DE, ES, FR} translation. Takes a 16 kHz mono WAV and produces a transcript. Not a streaming model; word/segment timestamps are upstream experimental and not exposed in the v1 port. Downloads Quantization Download Size WER (LibriSpeech test clean) : : F32 canary 1b flash F32.gguf 3.3 GB 1.62% F16 canary 1b flash F16.gguf 1.7 GB 1.62% Q8 0 canary 1b flash Q8 0.gguf 1.0 GB 1.62% Q6 K canary 1b flash Q6 K.gguf 818 MB 1.65% Q5 K M canary 1b flash Q5 K M.gguf 734 MB 1.64% Q4 K M canary 1b flash Q4 K M.gguf 646 MB 1.59% WER measured on the full LibriSpeech test clean split (2620 utterances) with greedy decoding and no external LM. F32 reference baseline: 1.62%. NVIDIA's self reported numbe…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy