Model Download Evaluation Results Model Architecture API Platform License Citation Paper Link 👁️ DeepSeek V2 Lite Chat FP8 This model was converted from https://huggingface.co/deepseek ai/DeepSeek V2 Lite Chat to Block FP8 quantization (128x128) to be used as testing for DeepSeek V3/R1. Dense's intermediate size was padded to be a multiple of 128. Below is the original model card. DeepSeek V2: A Strong, Economical, and Efficient Mixture of Experts Language Model 1. Introduction Last week, the release and buzz around DeepSeek V2 have ignited widespread interest in MLA (Multi head Latent Attention)! Many in the community suggested open sourcing a smaller MoE model for in depth research. And now DeepSeek V2 Lite comes out: 16B total params, 2.4B active params, scratch training with 5.7T tokens Outperforms 7B dense and 16B MoE on many English & Chinese benchmarks Deployable on single 40G GPU, fine tunable on 8x80G GPUs DeepSeek V2, a strong Mixture of Experts (MoE) language model characterized by economical training and efficient inference. DeepSeek V2 adopts innovative architectures including Multi head Latent Attention (MLA) and DeepSeekMoE. MLA guarantees efficient inference throug…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy