DeepSeek V4 Flash GGUF (community quants) Quantized variants of deepseek ai/DeepSeek V4 Flash (284B params, 13B active per token). Quants Quant Approx size Recommended? ~~IQ1 S~~ ~~~54 GB~~ Removed ~~IQ2 M~~ ~~~87 GB~~ Removed ~~IQ3 XS~~ ~~~109 GB~~ Removed Q2 K ~96 GB Q3 K M ~125 GB Q4 K M ~161 GB Provenance All variants derived from Preyazz/DeepSeek V4 Flash Q8 0 GGUF , which is itself a lossless conversion of the original FP8 safetensors. Compatibility Requires llama.cpp built from PR 22378 (nisparks's wip/deepseek v4 support branch) or later. The deepseek4 architecture is not yet in stable llama.cpp releases. For Strix Halo / consumer ROCm: build with GGML HIP NO VMM=ON (VMM=ON currently crashes on gfx1151 — see ROCm Issue 6146).
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy