📰 Tech Blog 1. Model Introduction Kimi K2 Thinking is the latest, most capable version of open source thinking model. Starting with Kimi K2, we built it as a thinking agent that reasons step by step while dynamically invoking tools. It sets a new state of the art on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks by dramatically scaling multi step reasoning depth and maintaining stable tool use across 200–300 sequential calls. At the same time, K2 Thinking is a native INT4 quantization model with 256k context window, achieving lossless reductions in inference latency and GPU memory usage. Key Features Deep Thinking & Tool Orchestration : End to end trained to interleave chain of thought reasoning with function calls, enabling autonomous research, coding, and writing workflows that last hundreds of steps without drift. Native INT4 Quantization : Quantization Aware Training (QAT) is employed in post training stage to achieve lossless 2x speed up in low latency mode. Stable Long Horizon Agency : Maintains coherent goal directed behavior across up to 200–300 consecutive tool invocations, surpassing prior models that degrade after 30–50 steps. 2. Model Summary…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy