canada quant/DeepSeek V4 Flash W4A16 FP8 Mixed precision quantization of deepseek ai/DeepSeek V4 Flash — W4A16 INT4 on routed experts + FP8 block 128×128 on attention — that loads cleanly on Hopper datacenter GPUs and on consumer grade Blackwell. Recipe topology mirrors RedHatAI/DeepSeek V4 Flash NVFP4 FP8 ; routed expert format is W4A16 (Marlin) instead of NVFP4 for compatibility with SM 9.x / SM 12.x kernels. TL;DR Recommended hardware 2× DGX Spark or 2× RTX PRO 6000, TP=2 Quality GSM8K 95.07–95.45% strict (8 shot); HumanEval pass@1 78.05–80.49% (strict, confirm run unsafe code ) Throughput 47–48 output tok/s @ bs=1 on RTX PRO 6000 TP=2 (TPOT 20.8 ms); 14–17 tok/s on DGX Spark TP=2 Differentiator Only quant of V4 Flash that serves on SM 9.x and SM 12.x; baseline for the W4A16 FP8 MTP successor Family / related artifacts Repo Role Relation to this artifact canada quant/DeepSeek V4 Flash W4A16 FP8 MTP successor Same recipe + BF16 MTP retained for 1.49× spec decode speedup at bs=1 canada quant/DeepSeek V4 Flash NVFP4 FP8 MTP sibling NVFP4 routed experts (Blackwell native), MTP retained canada quant/DeepSeek V4 Pro NVFP4 FP8 MTP larger sibling V4 Pro at NVFP4 with MTP, B300 only depl…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy