This is an Unsloth NVFP4 quantized checkpoint calibrated on the Hugging Face UltraChat dataset with sequences up to 16K context length and an approximately 2M token calibration budget. Qwen3.6 27B NVFP4 Runtime Benchmarks Full GSM8K and MMLU Pro evaluation was run with vLLM, chat template application, thinking disabled, and seed 3407 . These results are for the served NVFP4 runtime path. Model GSM8K MMLU Pro : : Qwen/Qwen3.6 27B 0.20 0.64 unsloth/Qwen3.6 27B NVFP4 0.20 0.63 vLLM Run Instructions Increase max model len only after checking available GPU memory. Multi Token Prediction (MTP) This checkpoint includes the MTP module, so it can act as its own speculative draft for faster decoding. vLLM: SGLang: [!Note] This repository contains model weights and configuration files for the post trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. Following the February release of the Qwen3.5 series, we're pleased to share the first open weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real world utility, offering developers a more intui…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy