DeepSeek V4: Towards Highly Efficient Million Token Context Intelligence Technical Report 👁️ Note: DeepSeek V4 Pro DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder. For more details, refer to: https://github.com/deepseek ai/DeepSpec Introduction We present a preview version of DeepSeek V4 series, including two strong Mixture of Experts (MoE) language models — DeepSeek V4 Pro with 1.6T parameters (49B activated) and DeepSeek V4 Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens . DeepSeek V4 series incorporate several key upgrades in architecture and optimization: 1. Hybrid Attention Architecture: We design a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to dramatically improve long context efficiency. In the 1M token context setting, DeepSeek V4 Pro requires only 27% of single token inference FLOPs and 10% of KV cache compared with DeepSeek V3.2. 2. Manifold Constrained Hyper Connections (mHC): We incorporate mHC to strengthen conventional residu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy