[ModelPage] : https://static.stepfun.com/blog/step 3.7 flash/ 1. Introduction GGUF quantizations of stepfun ai/Step 3.7 Flash . Step 3.7 Flash is a 198B parameter sparse Mixture of Experts vision language model from StepFun ai, activating ~11B parameters per token for up to 400 t/s throughput. It pairs a 196B parameter language backbone with a 1.8B parameter vision encoder for native image understanding, supports a 256K context window, and offers three selectable reasoning levels (low / medium / high) to balance speed, cost, and depth. Built for agentic workloads — tool calling, multi step reasoning, code, and math — with native multilingual coverage. A separate mmproj projector ships alongside the language quants for multimodal inference. With 128 GB of unified memory (Mac Studio, DGX Spark, Ryzen AI Max+ 395, etc.), you can privately host Step 3.7 Flash: Q4 quants and below run at full 256K context with high precision. 2. Files File Quant Size Notes : Step 3.7 flash BF16.gguf BF16 394 GB Full precision reference. Step 3.7 flash Q8 0.gguf Q8 0 209 GB Near lossless. Does not use imatrix. Step 3.7 flash Q4 K S.gguf Q4 K S 112 GB imatrix calibrated. Balanced quality / size. Step 3.7…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy