LLaDA2.0 mini preview LLaDA2.0 mini preview is a diffusion language model featuring a 16BA1B Mixture of Experts (MoE) architecture. As an enhanced, instruction tuned iteration of the LLaDA series, it is optimized for practical applications. Benchmark Ling mini 2.0 LLaDA MoE 7B A1B Instruct LLaDA2.0 mini preview : : : : : : : : Average 74.60 59.72 66.89 Knowledge MMLU 82.15 67.18 72.49 MMLU PRO 63.72 44.64 49.22 CMMLU 80.84 64.30 67.53 C EVAL 82.10 63.93 66.54 Reasoning squad2.0 75.56 86.81 85.61 drop 78.80 79.77 79.49 korbench 62.72 38.40 37.26 Coding CruxEval O 76.12 42.38 61.88 mbpp 84.07 70.02 77.75 MultiPL E 67.09 52.53 62.43 humaneval 85.98 61.59 80.49 Bigcodebench Full 35.00 20.44 30.44 Math GSM8K 94.62 82.41 89.01 math 94.66 58.68 73.50 Agent & Alignment BFCL Live 53.98 63.09 74.11 IFEval strict prompt 76.16 59.33 62.50 🚀 Performance Highlights + Leading MoE Architecture : The open source Mixture of Experts (MoE) diffusion large language model continually trained on the Ling2.0 series with approximately 20 trillion tokens . + Efficient Inference : With 16 billion total parameters , only 1.4 billion are activated during inference. LLaDA2.0 mini preview significantly reduces…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy