Voxtral Mini 4B Realtime 2602: transcribe.cpp GGUF GGUF conversions of mistralai/Voxtral Mini 4B Realtime 2602 for use with transcribe.cpp. Ported from upstream commit 2769294, pinned 2026 06 06. Validated against the Transformers reference at transcribe.cpp commit 483c122 on 2026 06 06. Streaming audio LLM speech to text. A ~970M causal audio encoder (left pad causal conv stem + 32 layer sliding window RoPE transformer) feeds a 4 frame group projector whose audio embeddings are added onto a ~3.4B Ministral decoder (26 layers, GQA 32/8, NEOX RoPE) with delay token latency conditioning, emitting one text token per 80 ms audio slot (12.5 Hz). Takes a 16 kHz mono WAV and supports both incremental streaming (configurable latency/quality via stream chunk ms and stream voxtral delay) and offline transcription with a byte equal final transcript. Architecturally distinct from the offline Voxtral 2507 family — own arch, streaming frontend, causal encoder, additive audio fusion. Downloads Quantization Download Size WER (LibriSpeech test clean) : : BF16 Voxtral Mini 4B Realtime 2602 BF16.gguf 8.87 GB 2.08% F16 Voxtral Mini 4B Realtime 2602 F16.gguf 8.88 GB 2.09% Q8 0 Voxtral Mini 4B Realtime…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy