📰 Step3 Model Blog 📄 Step3 System Blog Introduction Step3 is our cutting edge multimodal reasoning model—built on a Mixture of Experts architecture with 321B total parameters and 38B active. It is designed end to end to minimize decoding costs while delivering top tier performance in vision–language reasoning. Through the co design of Multi Matrix Factorization Attention (MFA) and Attention FFN Disaggregation (AFD), Step3 maintains exceptional efficiency across both flagship and low end accelerators. Step3 model card: Config Value Number of Layers (Dense layer included) 61 Number of Dense Layers 5 Hidden Dimension 7168 Attention Mechanism MFA Low rank Query Dimension 2048 Number of Query Heads 64 Head Dimension 256 Number of Experts 48 Selected Experts per Token 3 Number of Shared Experts 1 Max Context Length 65536 Tokenizer Deepseek V3 Total Parameters (LLM) 316B Activated Params per Token 38B Total Parameters (VLM) 321B Evaluation Results Deployment [!Note] Step3's API is accessible at https://platform.stepfun.com/, where we offer OpenAI compatible API for you. Inference with Hugging Face Transformers We introduce ho…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy