canada quant/DeepSeek V4 Flash W4A16 FP8 MTP W4A16 INT4 routed experts + FP8 block 128×128 attention + BF16 Multi Token Prediction (MTP) draft head retained — the first DeepSeek V4 Flash quantization that ships a working MTP block, giving ~1.5× speculative decoding (spec decode) speedup at bs=1 with no quality cost. Extends the W4A16 FP8 predecessor by patching the transformers calibration path so the MTP block survives the load. TL;DR Recommended hardware RTX PRO 6000 Blackwell at TP=2 (2 GPUs/replica) or TP=4 (4 GPUs/replica) — both validated · or 8× H200 TP=2 Quality GSM8K 93.71% (8 shot strict); HumanEval 84.76% pass@1; MMLU 86.88% Throughput RTX PRO 6000 98.83 @ TP=2 / 107.32 @ TP=4 at bs=1; 88.35 on H200 TP=2 MTP acceptance 89% calibrated workload / 70% on random prompts at bs=1 k=1 Spec decode speedup 1.49× at bs=1, k=1 (TPOT 6.02 ms vs 8.93 ms, same artifact) Differentiator First V4 Flash W4A16 quant where MTP survives the calibration load; transformers 5.8.1 silently strips MTP keys by default Family / related artifacts Repo Role Relation to this artifact canada quant/DeepSeek V4 Flash W4A16 FP8 predecessor Same W4A16 + FP8 recipe; MTP dropped at load (the bug this artifac…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy