Qwen3 Coder 30B A3B DFlash Paper GitHub Blog DFlash is a novel speculative decoding method that utilizes a lightweight block diffusion model for drafting. It enables efficient, high quality parallel drafting that pushes the limits of inference speed. This model is the drafter component. It must be used in conjunction with the target model Qwen/Qwen3 Coder 30B A3B Instruct . 📊 Training Data & Efficiency Qwen3 Coder 30B A3B DFlash is trained on 289K samples , composed of: Code split from nvidia/Nemotron Post Training Dataset v2 theblackcat102/evol codealpaca v1 Approximately 2.8K Cline execution traces collected by ourselves Despite being trained on significantly less data , DFlash already outperforms EAGLE 3 in inference acceleration. In comparison, lmsys/SGLang EAGLE3 Qwen3 Coder 30B A3B Instruct SpecForge is trained on the open perfect blend dataset with 1.4M samples , nearly 5× more data than DFlash. This result highlights the training efficiency and scalability of DFlash, and suggests that further scaling the training data can unlock even greater acceleration gains. 🚀 Quick Start SGLang Installation Launch Server Usage vLLM Installation Launch Server Usage Transformers Install…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy