Model Download Evaluation Results Model Architecture API Platform License Citation Paper Link 👁️ DeepSeek V2: A Strong, Economical, and Efficient Mixture of Experts Language Model 1. Introduction Today, we’re introducing DeepSeek V2, a strong Mixture of Experts (MoE) language model characterized by economical training and efficient inference. It comprises 236B total parameters, of which 21B are activated for each token. Compared with DeepSeek 67B, DeepSeek V2 achieves stronger performance, and meanwhile saves 42.5% of training costs, reduces the KV cache by 93.3%, and boosts the maximum generation throughput to 5.76 times. We pretrained DeepSeek V2 on a diverse and high quality corpus comprising 8.1 trillion tokens. This comprehensive pretraining was followed by a process of Supervised Fine Tuning (SFT) and Reinforcement Learning (RL) to fully unleash the model's capabilities. The evaluation results validate the effectiveness of our approach as DeepSeek V2 achieves remarkable performance on both standard benchmarks and open ended generation evaluation. 2. Model Downloads Model Context Length Download : : : : : : DeepSeek V2 128k 🤗 HuggingFace DeepSeek V2 Chat (RL) 128k 🤗 Hugging…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy