๐ฐ Tech Blog ๐ Paper 1. Model Introduction Kimi K2 Instruct 0905 is the latest, most capable version of Kimi K2. It is a state of the art mixture of experts (MoE) language model, featuring 32 billion activated parameters and a total of 1 trillion parameters. Key Features Enhanced agentic coding intelligence: Kimi K2 Instruct 0905 demonstrates significant improvements in performance on public benchmarks and real world coding agent tasks. Improved frontend coding experience: Kimi K2 Instruct 0905 offers advancements in both the aesthetics and practicality of frontend programming. Extended context length: Kimi K2 Instruct 0905โs context window has been increased from 128k to 256k tokens, providing better support for long horizon tasks. 2. Model Summary : : : : Architecture Mixture of Experts (MoE) Total Parameters 1T Activated Parameters 32B Number of Layers (Dense layer included) 61 Number of Dense Layers 1 Attention Hidden Dimension 7168 MoE Hidden Dimension (per Expert) 2048 Number of Attention Heads 64 Number of Experts 384 Selected Experts per Token 8 Number of Shared Experts 1 Vocabulary Size 160K Context Length 256Kโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy