Voxtral Mini 3B 2507: transcribe.cpp GGUF GGUF conversions of mistralai/Voxtral Mini 3B 2507 for use with transcribe.cpp. Ported from upstream commit 3060fe3, pinned 2026 06 06. Validated against the Transformers reference at transcribe.cpp commit 483c122 on 2026 06 06. Offline audio LLM speech to text and speech translation. A Whisper large v3 bidirectional audio encoder feeds a 4 frame group projector (375 audio tokens per 30 s chunk) into a Ministral 3B causal LM (30 layers, GQA 32/8, NEOX RoPE, SwiGLU) via audio token injection. Takes a 16 kHz mono WAV and produces a transcript via greedy decoding; speech translation runs through the mistral common instruct template. The smaller sibling of Voxtral Small 24B — same encoder, projector, log mel frontend, and tekken tokenizer, with a 3B decoder in place of Mistral Small 24B. Downloads Quantization Download Size WER (LibriSpeech test clean) : : BF16 Voxtral Mini 3B 2507 BF16.gguf 9.37 GB 1.88% F16 Voxtral Mini 3B 2507 F16.gguf 9.38 GB 1.89% Q8 0 Voxtral Mini 3B 2507 Q8 0.gguf 5.00 GB 1.87% Q6 K Voxtral Mini 3B 2507 Q6 K.gguf 3.87 GB 1.87% Q5 K M Voxtral Mini 3B 2507 Q5 K M.gguf 3.46 GB 1.91% Q4 K M Voxtral Mini 3B 2507 Q4 K M.gguf 2…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy