Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants. KLD Benchmarks: Note MXFP4 MOE is known to have worse KLD vs disk space use UD Q4 K XL instead. mmproj / vision works as well [ModelPage] : https://static.stepfun.com/blog/step 3.7 flash/ 1. Introduction Step 3.7 Flash is a 198B parameter sparse Mixture of Experts (MoE) vision language model that combines a 196B parameter language backbone with a 1.8B parameter vision encoder for native image understanding. Engineered for high frequency production workloads, it activates approximately 11B parameters per token and delivers a throughput of up to 400 tokens per second. Step 3.7 Flash supports a 256k context window and offers three selectable reasoning levels (low, medium, and high) so developers can easily balance speed, cost, and cognitive depth. We built Step 3.7 Flash for developers who need to scale agentic workflows that combine perception, search, and reasoning. It is designed to handle intensive tasks such as parsing massive financial reports in one pass, running multi step search loops with cross source verification, or operating concurrent coding agents in high throughput pipelines. 2. Capabili…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy