🚀 Qwen36 27B GPTQ Pro 4Bit Welcome to Qwen36 27B GPTQ Pro 4Bit – a titan of reasoning and generation, elegantly squeezed into a remarkably efficient 4 bit package. It punches leagues above its weight class while keeping your VRAM happy and your inference speeds blazingly fast! Thank you Qwen team for another amazing model. 🌟 Why the "Pro"? This isn't your average quantization. We used the GPTQ Pro framework combined with the FOEM (First Order Error Metric) approach. This advanced technique carefully preserves the most critical weights during the 4 bit compression process by evaluating the exact impact of quantization on the model's loss landscape. The result? Near Lossless Performance : Enjoy the profound reasoning, coding prowess, and vast knowledge of a 27 Billion parameter model, but with a drastically reduced memory footprint. Marlin Optimized : Ready out of the box for Marlin kernels to deliver maximum token per second throughput in serving engines like vLLM. Consumer Hardware Friendly : Fit a massive 27B powerhouse model on consumer GPUs with room to spare for massive context lengths! This repository contains a 4 bit GPTQ Pro quantization of unsloth/Qwen3.6 27B , produced w…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy