MiniMax M2.5 (Mixed Precision BF16 + INT4 AWQ) Changelog 2026 02 15: Requant to ensure quality batch size=1 (thread) and addition of Greek language and 34 multilingual dataset (request) 2026 02 14: Original quant, using LLMcompressor's new feature batch size=32 . batch size may negatively impact calibration due to truncating or padding datasets, defeating the careful selection I made. Overview This strives to be the highest quality quant that can run on 192GiB VRAM [!TIP] 💡This is a sister model to mratsim/MiniMax M2.5 FP8 INT4 AWQ with the original model FP8 weights pre dequantized to BF16. This makes it compatible with 8x3090 systems (which don't have hardware FP8) and also compatible with SGLang for an extra 3 GiB in VRAM. It features: That model has ensured that all experts are calibrated, not doing so is extremely detrimental, PR: https://github.com/vllm project/llm compressor/pull/2171 💡 [Click me!] Visual showcase of why ensuring quantization of all MoE experts is important Source: https://avtc.github.io/aquarium side by side/ Context: https://github.com/ModelCloud/GPTQModel/pull/2235 Mixed precision with: self attention weights dequantized from the official version. exper…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy