We worked with Mistral to fix Mistral Medium 3.5 inference issues affecting some implementations ( not related to Unsloth or our quants). The issue came from a YaRN parsing quirk in implementations like transformers and llama.cpp. Setting mscale all dim from 1 to 0 fixes it, including the model forgetting previous conversations. Mistral has now pushed these fixes to their official repo. Read our How to Run Mistral 3.5 Guide! See Unsloth Dynamic 2.0 GGUFs for our quantization benchmarks. Mistral Medium 3.5 128B Mistral Medium 3.5 is our first flagship merged model. It is a dense 128B model with a 256k context window, handling instruction following, reasoning, and coding in a single set of weights. Mistral Medium 3.5 replaces its predecessor Mistral Medium 3.1 and Magistral in Le Chat. It also replaces Devstral 2 in our coding agent Vibe. Concretely, expect better performance for instruct, reasoning and coding tasks in a new unified model in comparison with our previous released models. Reasoning effort is configurable per request, so the same model can answer a quick chat reply or work through a complex agentic run. We trained the vision encoder from scratch to handle variable image…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy