Mistral Voxtral Mini 4B Realtime INT4 NF4 Submission This repository contains an INT4 NF4 quantized version of: mistralai/Voxtral Mini 4B Realtime 2602 Compression technique Base model: mistralai/Voxtral Mini 4B Realtime 2602 Quantization method: BitsAndBytes 4 bit NF4 Double quantization: enabled Compute dtype: BF16 on A100, otherwise FP16 Architecture changes: none Distillation: none Fine tuning: none Challenge alignment The model remains based on Voxtral Realtime. The compression approach focuses on weight quantization. The challenge evaluation is expected to focus on ASR quality using WER, followed by energy efficiency ranking among qualifying submissions. Local validation performed A tiny FLEURS smoke test was run on three languages: English: en us French: fr fr Hindi: hi in The smoke test used one sample per language and was intended only to verify that the quantized model loads and transcribes. Observed smoke test macro WER: BF16 baseline: 0.668129 INT4 NF4: 0.650585 Because this test used only one sample per language, these numbers should not be interpreted as a final benchmark. They only indicate that the INT4 checkpoint is functional and not obviously broken. Serving Inte…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy