1. Model Introduction JoyAI LLM Flash is a state of the art medium sized instruct language model with 3 billion activated parameters and 48 billion total parameters. JoyAI LLM Flash was pretrained on 20 trillion text tokens using Muon optimizer, followed by large scale supervised fine tuning (SFT), direct preference optimization (DPO), and reinforcement learning (RL) across diverse environments. JoyAI LLM Flash achieves strong performance across frontier knowledge, reasoning, coding tasks and agentic capabilities. Key Features Fibration Policy Optimization: Introduces fiber bundle theory into reinforcement learning, proposing a novel optimization framework, FiberPO. This method is specifically designed to handle the challenges of large scale and heterogeneous agent training, improving stability and robustness under complex data distributions. paper link Training Inference Collaboration: apply Muon optimizer with dense MTP, develop novel optimization techniques to resolve instabilities while scaling up, delivering 1.3× to 1.7× the throughput of the non MTP version. Agentic Intelligence: designed for tool use, reasoning, and autonomous problem solving. 2. Model Summary : : : : Archit…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy