LongCat Flash Lite Tech Report 📄 Model Introduction We introduce LongCat Flash Lite, a non thinking 68.5B parameter Mixture of Experts (MoE) model with approximately 3B activated parameters, supporting a 256k context length through the YaRN method. Building upon the LongCat Flash architecture, LongCat Flash Lite distinguishes itself through the integration of an N gram embedding table designed to enhance both model performance and inference speed. Despite allocating over 30B parameters to embeddings, LongCat Flash Lite not only outperforms parameter equivalent MoE baselines but also demonstrates exceptional competitiveness against existing models of comparable scale, particularly in the agentic and coding domains. Key Features 🌟 Superior Scaling Efficiency: A Better Alternative to MoE Through comprehensive scaling experiments across diverse scenarios, we identify specific regimes where embedding scaling achieves a superior Pareto frontier compared to increasing the number of experts, thereby offering a highly efficient alternative for model scaling. We further delineate a comprehensive set of architectural factors that determine embedding scaling efficacy, encompassing inte…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy