Mistral Large 3 675B Instruct 2512 NVFP4 From our family of large models, Mistral Large 3 is a state of the art general purpose Multimodal granular Mixture of Experts model with 41B active parameters and 675B total parameters trained from the ground up with 3000 H200s. This model is the instruct post trained version, fine tuned for instruction tasks, making it ideal for chat, agentic and instruction based use cases. Designed for reliability and long context comprehension It is engineered for production grade assistants, retrieval augmented systems, scientific workloads, and complex enterprise workflows. [!Tip] This checkpoint in particular is a post training activation quantized version of Mistral Large 3 675B Instruct 2512. It was created using llm compressor as part of a collaboration with teams from vLLM & Red Hat. Special thanks goes out to Dipika Sikka, Kyle Sayers, Eldar Kurtić, and Tyler Michael Smith. Mistral Large 3 is deployable on premises in: FP8 on a single node of B200s or H200s. NVFP4 on a single node of H100s or A100s. Key Features Mistral Large 3 consists of two main architectural components: A Granular MoE Language Model with 673B params and 39B active A 2.5B Visi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy