LLaMA3.1 8B Instruct DFlash UltraChat Paper GitHub Blog DFlash is a novel speculative decoding method that utilizes a lightweight block diffusion model for drafting. It enables efficient, high quality parallel drafting that pushes the limits of inference speed. This model is the drafter component. It must be used in conjunction with the target model meta llama/Llama 3.1 8B Instruct . 📊 Training Data LLaMA3.1 8B Instruct DFlash UltraChat is trained on Ultrachat 200K and ShareGPT datasets, aiming to align with EAGLE 3 training data. The assistant reponses in the datasets are regenerated by meta llama/Llama 3.1 8B Instruct . 🚀 Quick Start SGLang Installation Launch Server Usage vLLM Installation Launch Server Usage Transformers Installation Inference Evaluation DFlash consistently achieves higher speedups than the state of the art speculative decoding method EAGLE 3 . All experiments are conducted using SGLang on a single B200 GPU . For EAGLE 3, we evaluate two speculative decoding configurations: speculative num steps 7 , speculative eagle topk 10 , speculative num draft tokens 10 speculative num steps 7 , speculative eagle topk 10 , speculative num draft tokens 60 , which is the o…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy