This model is obtained similarly to how the RedHatAI/Qwen3.6 35B A3B NVFP4 was obtained with the following compression script using llm compressor . NOTE : Unlike the aforementioned model, the linear attn layers have been quantized as well in this model to save memory for longer context lengths on RTX 5090 GPUs. Click the dropdown to see the full quantization script.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy