Mistral Medium 3.5 128B — NVFP4 (W4A4 FP4), v3 calibration NVFP4 (W4A4 FP4) quantization of mistralai/Mistral Medium 3.5 128B , produced with llm compressor and saved in the compressed tensors nvfp4 pack quantized format. The vision tower, multi modal projector, embeddings, and lm head are kept at bf16 — the language model body (88 Ministral3DecoderLayer blocks) is the only thing in FP4. Designed to be served by vLLM on Blackwell class GPUs (SM 12.0+ — RTX 50 series / B series) using the FlashInfer NVFP4 GEMM kernel. This release replaces the prior NVFP4 build in this repo. Two things changed: 1. Calibration corpus expanded from 20 samples on a 4 source mix to 256 samples drawn from a 2560 sample mixed format corpus that exercises Mistral chat templated, Anthropic XML, and OpenAI tool JSON surfaces. 2. YARN config corrected. The bf16 source had rope yarn log multiplier=1.0 upstream, which broke long context generation. This build's config.json has the corrected 0.0 value, matching Mistral's config fix commit. Compression summary Component Tensors Dtype Approx size language model (88 decoder layers) 2641 NVFP4 (W4A4 FP4 + scales) 63.81 GiB vision tower (Pixtral) 434 bf16 4.65 GiB mu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy