Step 3.5 Flash Step 3.5 Flash 1. Introduction Step 3.5 Flash (visit website) is our most capable open source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token. This "intelligence density" allows it to rival the reasoning depth of top tier proprietary models, while maintaining the agility required for real time interaction. 2. Key Capabilities Deep Reasoning at Speed : While chatbots are built for reading, agents must reason fast. Powered by 3 way Multi Token Prediction (MTP 3), Step 3.5 Flash achieves a generation throughput of 100–300 tok/s in typical usage (peaking at 350 tok/s for single stream coding tasks). This allows for complex, multi step reasoning chains with immediate responsiveness. A Robust Engine for Coding & Agents : Step 3.5 Flash is purpose built for agentic tasks, integrating a scalable RL framework that drives consistent self improvement. It achieves 74.4% on SWE bench Verified and 51.0% on Terminal Bench 2.0 , proving its ability to handle sophisticated, long horizon tasks with…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy