LongCat Flash Chat Model Introduction We introduce LongCat Flash, a powerful and efficient language model with 560 billion total parameters, featuring an innovative Mixture of Experts (MoE) architecture. The model incorporates a dynamic computation mechanism that activates 18.6B∼31.3B parameters (averaging∼27B) based on contextual demands, optimizing both computational efficiency and performance. To achieve advanced training and inference efficiency, we employ a shortcut connected architecture that expands computation communication overlap window, achieving over 100 tokens per second (TPS) for inference cost effectively. Our comprehensive training and scaling strategies ensure stable, efficient training, while tailored data strategies enhance model performance. Now we release LongCat Flash Chat, a non thinking foundation model that delivers highly competitive performance among leading models, with exceptional strengths in agentic tasks. Key Features 🌟 Scalable Architectural Design for Computational Efficiency LongCat Flash is designed and optimized under two key principles: efficient computation utilization, as well as efficient training and inference. Specifically, (1) As not all…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy