[!TIP] Support this work → · X · GitHub · REAP paper · Cerebras REAP DeepSeek V4 Flash 180B REAP pruned deepseek ai/DeepSeek V4 Flash. At a glance Base model deepseek ai/DeepSeek V4 Flash Format BF16 Total params 180B Active / token — Experts / layer 160 Layers 43 Hidden size 4096 Context 1,048,576 On disk size 103 GB Which variant should I pick? Variant Format Link DeepSeek V4 Flash 162B BF16 link DeepSeek V4 Flash 162B GGUF GGUF link DeepSeek V4 Flash 180B (this) BF16 link DeepSeek V4 Flash 180B GGUF GGUF link DeepSeek V4 Flash 213B BF16 link 180B parameters K160 REAP pruned 200K context MTP speculative decoding This is a pruned and quantized DeepSeek V4 Flash that runs on a single DGX Spark. It is not the original model. It is a derivative built to fit into 128 GB of host memory while keeping the full 200,000 token context window alive. The goal was simple: take one of the best open reasoning models available and make it runnable on a desktop AI workstation without losing what makes it useful. The result serves at about 24 tok/s decode with 2 token speculative decoding, and it retains a needle buried at 200K context. What this is Base: deepseek ai/DeepSeek V4 Flash Pruning: REAP…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy