Voxtral Small 24B 2507: transcribe.cpp GGUF GGUF conversions of mistralai/Voxtral Small 24B 2507 for use with transcribe.cpp. Ported from upstream commit da5b424, pinned 2026 06 05. Validated against the Transformers reference at transcribe.cpp commit dac22fa on 2026 06 05. Offline audio LLM speech to text and speech translation. A Whisper large v3 bidirectional audio encoder feeds a 4 frame group projector (375 audio tokens per 30 s chunk) into a Mistral Small 24B causal LM (40 layers, GQA 32/8, NEOX RoPE, SwiGLU) via audio token injection. Takes a 16 kHz mono WAV and produces a transcript via greedy decoding. The larger sibling of Voxtral Mini 3B — same encoder, projector, frontend, and tokenizer, with a scaled up decoder. Downloads Quantization Download Size WER (LibriSpeech test clean) : : BF16 Voxtral Small 24B 2507 BF16.gguf 48.54 GB 1.56% F16 Voxtral Small 24B 2507 F16.gguf 48.55 GB 1.57% Q8 0 Voxtral Small 24B 2507 Q8 0.gguf 25.81 GB 1.56% Q6 K Voxtral Small 24B 2507 Q6 K.gguf 19.94 GB 1.58% Q5 K M Voxtral Small 24B 2507 Q5 K M.gguf 17.14 GB 1.60% Q4 K M Voxtral Small 24B 2507 Q4 K M.gguf 14.30 GB 2.11% WER measured on the full LibriSpeech test clean split (2620 utterances)…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy