Mistral Medium 3.5 128B — GGUF (imatrix v3) GGUF quants of mistralai/Mistral Medium 3.5 128B for use with llama.cpp (and tools downstream of it). Calibrated with a custom 1552 chunk importance matrix and shipped alongside a separate mmproj for multimodal use. [!Important] The upstream YARN scaling config originally had rope yarn log multiplier=1.0 , which broke long context generation. The bf16 source GGUF used for these quants has rope yarn log multiplier=0.0 (the corrected value, matching Mistral's config fix commit). Older GGUFs in the wild that were converted before that fix may still produce degraded outputs on long contexts. Files File Quant Size (approx) Notes Mistral Medium 3.5 128B Q4 K M v3.gguf Q4 K M 70 GB 4.79 BPW, default 4 bit pick Mistral Medium 3.5 128B Q5 K M v3.gguf Q5 K M ~83 GB 5.69 BPW Mistral Medium 3.5 128B Q6 K v3.gguf Q6 K ~96 GB 6.56 BPW, near lossless Mistral Medium 3.5 128B mmproj bf16.gguf bf16 5.1 GB Vision tower + multi modal projector. Required for image input. Pair with any of the LM GGUFs above via mmproj . An earlier upload of Q4 K M.gguf and Q6 K.gguf (built with a 91 chunk imatrix, before the YARN config was fully validated end to end) has been…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy