Qwen3.6 27B DFlash Paper GitHub Blog This model is still under training, and inference engine support may not be fully available yet due to architectural changes, including causal SWA layers. DFlash is a novel speculative decoding method that utilizes a lightweight block diffusion model for drafting. It enables efficient, high quality parallel drafting that pushes the limits of inference speed. This model is the drafter component. It must be used in conjunction with the target model Qwen/Qwen3.6 27B . Quick Start Installation vLLM (We temporarily modify the installation through this PR to support interleaved SWA and ensure correct handling of target hidden states for optimal performance): SGLang: Launch Server vLLM: SGLang: Usage Benchmark Results N/A Acknowledgements Special thanks to David Wang for his outstanding engineering support on this project. We are also grateful to Modal, InnoMatrix, and Yotta Labs for providing the compute resources used to train this draft model. Citation If you find DFlash useful, please cite our work. To share feedback on DFlash or request new model support, please fill out this form: DFlash Feedback.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy