Step 3.7 Flash GGUF My custom IQ4 XS GGUF quantization for stepfun ai/Step 3.7 Flash I've also modified the chat template also adds a preserve thinking option, which preserves thinking across user turns and can improve the experience when prompt processing speed is a bottleneck. Quant Recipes Recipe Quant Size Default type Tensor specific overrides IQ4 XS 101784.88 MiB (4.34 BPW) Q6 K ffn down exps=iq4 xs , ffn gate exps=iq4 xs , ffn up exps=iq4 xs Related Files File Description Step 3.7 Flash MTP Q8 0.gguf Q8 0 MTP weights Step 3.7 Flash mmproj BF16.gguf BF16 multimodal projector Step 3.7 Flash mmproj F16.gguf F16 multimodal projector Step 3.7 Flash mmproj Q8 0.gguf Q8 0 multimodal projector Usage Here's an example script:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy