LLaDA MoE LLaDA MoE is a new and upgraded series of the LLaDA diffusion language model. This pre release includes two cutting edge models: LLaDA MoE 7B A1B Base : A base pre trained model designed for research and secondary development. LLaDA MoE 7B A1B Instruct : An instruction tuned model optimized for practical applications. LLaDA MoE 7B A1B Instruct TD : A specialized instruction tuned model, further optimized for accelerated inference using Trajectory Distillation. 🚀 Performance Highlights Leading MoE Architecture : The first open source Mixture of Experts (MoE) diffusion large language model , pre trained from scratch on approximately 20 trillion tokens . Efficient Inference : With 7 billion total parameters , only 1.4 billion are activated during inference. LLaDA MoE significantly reduces computational costs while outperforming open source dense models of similar scale. Impressive Performance on Code & Complex Reasoning : Excels in tasks such as code generation and advanced mathematical reasoning , demonstrating strong reasoning capabilities. Tool Use : Supports tool calling and achieves excellent performance in complex agent based tasks. Open & Extensible : Fully open sour…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy