MiniMax M2.7 AWQ 4bit (W4A16) W4A16 quantization of MiniMaxAI/MiniMax M2.7 , produced with llm compressor . Format: compressed tensors pack quantized, int4 weights / fp16 activations Group size: 128, symmetric Calibration: data free, MSE observer Kept in BF16: MoE routing gates and lm head only — every other Linear is quantized, matching the ignore list from cyankiwi/MiniMax M2.5 AWQ 4bit . vLLM
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy