SDAR Arxiv • 💻Github Repo • 🤗Model Collections Introduction SDAR ( S ynergy of D iffusion and A uto R egression) model is a new large language model that integrates autoregressive (AR) and discrete diffusion modeling strategies. It combines the efficient training paradigm of AR models with the highly parallel inference capability of diffusion models, while delivering performance fully on par with SOTA open source AR models. At the same time, SDAR sets a new benchmark as the most powerful diffusion language model to date. We highlight three major conclusions from our study: [!IMPORTANT] Take home message Balanced Efficiency: SDAR unifies the efficient training of AR models with the parallel inference of diffusion, achieving both fast training and inference. Fair Comparisons: In rigorously controlled experiments, SDAR achieves on par general task performance with strong AR baselines, ensuring credibility and reproducibility. Superior Learning Efficiency: On complex scientific reasoning tasks (e.g., GPQA, ChemBench, Physics), SDAR shows clear gains over AR models of the same scale, approaching or even exceeding leading closed source systems. Inference Using the tailored infer…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy