ghecko78/Voxtral Mini 4B Realtime 2602 FP8 Dynamic FP8 quantized version of mistralai/Voxtral Mini 4B Realtime 2602 for faster inference and reduced memory usage. Overview Property Value Base Model mistralai/Voxtral Mini 4B Realtime 2602 Quantization FP8 Dynamic ( FP8 DYNAMIC ) Weight Quantization Symmetric, static, per channel → FP8 (E4M3) Activation Quantization Symmetric, dynamic, per token → FP8 (E4M3) Format compressed tensors (vLLM native) Quantized Size ~5.43 GB Tool llm compressor Date 2026 05 22 What is this? This is an FP8 quantized version of Mistral AI's Voxtral Mini 4B Realtime — a multilingual, streaming speech to text model. The quantization reduces: Memory footprint by ~50% (from ~8 GB to ~4 GB) Inference latency through hardware accelerated FP8 tensor operations Time to first token with smaller weight transfers All while maintaining near identical transcription quality to the original BF16 model. Supported Languages English, French, German, Spanish, Italian, Portuguese, Dutch, Polish, Swedish, Danish, Finnish, Norwegian (Bokmål), Hindi Quantization Details The quantization was performed using llm compressor with the FP8 DYNAMIC scheme: Weights : Quantized with symm…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy