Mistral Medium 3.5 128B NVFP4 2026 05 03 Config Fix: rope scaling.mscale all dim in config.json has been corrected from 1.0 to 0.0 , matching the upstream Mistral fix. This parameter only affects YaRN attention scaling at inference time and has no impact on the quantized weights (this model uses data free NVFP4A16 weight only quantization — no recalibration needed). If you have already downloaded this model, simply update this field in config.json . No need to re download the weight files. NVFP4 (W4A16) quantization of mistralai/Mistral Medium 3.5 128B, created with llm compressor in compressed tensors nvfp4 pack quantized format. Designed for vLLM inference on Blackwell class GPUs (compute capability 10.0+, e.g. RTX 5090 / B200). Quantization Details Scheme: NVFP4A16 (weights quantized to NVIDIA FP4, activations unquantized / weight only) Algorithm: RTN (Round To Nearest) via QuantizationModifier with scheme="NVFP4A16" Calibration data: None. Fully data free quantization, no calibration samples used. Source model: mistralai/Mistral Medium 3.5 128B (FP8, per tensor, activation scheme: static ) Conversion: FP8 weights dequantized to BF16 ( weight weight scale inv ), then quantized t…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy