Qwen3.6 27B Claude Opus Sonnet Distilled NVFP4 MTP Claude Opus + Sonnet distilled Qwen 3.6 27B optimized for higher token efficiency, preserved deep reasoning, and MTP accelerated vLLM deployment on Blackwell class GPUs. This release is designed around one practical goal: make high quality local deployment feel fast enough, responsive enough, and efficient enough to use every day. Highlights Higher token efficiency : less budget wasted on invisible long form reasoning, more budget converted into visible answers Reasoning depth preserved : no obvious degradation in math, logic, or complex systems prompts in local spot checks MTP actually works : speculative decoding is not just included in the files, it shows measurable benefit in the tested deployment stack Better interactive UX : faster visible output, lower waiting time, and a much more usable local API experience In short: this is not about making the model think less, it is about making it spend more of its token budget on answers users can actually see. Why This Release The value of this repository is not just "an NVFP4 checkpoint". It is the combination of: Claude Opus + Sonnet distilled response style and reasoning organizat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy