Model Download Evaluation Results Model Architecture API Platform License Citation Paper Link 👁️ DeepSeek V2: A Strong, Economical, and Efficient Mixture of Experts Language Model 1. Introduction Last week, the release and buzz around DeepSeek V2 have ignited widespread interest in MLA (Multi head Latent Attention)! Many in the community suggested open sourcing a smaller MoE model for in depth research. And now DeepSeek V2 Lite comes out: 16B total params, 2.4B active params, scratch training with 5.7T tokens Outperforms 7B dense and 16B MoE on many English & Chinese benchmarks Deployable on single 40G GPU, fine tunable on 8x80G GPUs DeepSeek V2, a strong Mixture of Experts (MoE) language model characterized by economical training and efficient inference. DeepSeek V2 adopts innovative architectures including Multi head Latent Attention (MLA) and DeepSeekMoE. MLA guarantees efficient inference through significantly compressing the Key Value (KV) cache into a latent vector, while DeepSeekMoE enables training strong models at an economical cost through sparse computation. 2. News 2024.05.16: We released the DeepSeek V2 Lite. 2024.05.06: We released the DeepSeek V2. 3. Model Downloads…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy