LongCat Flash Thinking 2601 Tech Report 📄 Model Introduction We introduce an updated version of LongCat Flash Thinking, a powerful and efficient Large Reasoning Model (LRM) with 560 billion total parameters, built upon an innovative Mixture of Experts (MoE) architecture. Beyond inheriting the domain parallel training recipe in our previous version and maintaining highly competitive performance on traditional reasoning benchmarks, this update systematically strengthens agentic thinking capability through a carefully designed pipeline that combines environment scaling and subsequent task synthesis, followed by reliable and efficient large scale and multi environment reinforcement learning. To better adapt to the noise and uncertainty inherent in real world agentic tasks, we conduct systematic analysis and curriculum training over multiple types and levels of environmental noise, enabling robust performance under imperfect conditions. As a result, LongCat Flash Thinking 2601 achieves not only top tier benchmark performance in agentic tool use, agentic search, and tool integrated reasoning, but also substantially improved generalization in arbitrary out of distribution real worl…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy