DeepSeek V4 Flash JANGTQ2 DeepSeek V4 Flash — 79.6 GB on disk (down from 149 GB FP4+FP8 source) — uniform 2 bit JANGTQ quantization on routed experts + 8 bit affine on everything else + preserved MTP head. Source: deepseek ai/DeepSeek V4 Flash (43 transformer layers + 1 MTP head, 256 routed experts top 6 + 1 shared expert , 3 hash layers , MLA + mHC residuals, ~284 B total) Quantization: uniform 2 bit MXTQ on routed expert MLP + 8 bit affine on attention ( wq a/wq b/wkv/wo a/wo b ) / shared expert / Compressor / Indexer / embed / lm head / MTP. RMSNorms, router gate, mHC fn matrices, attn sink, ape stay fp16/fp32 passthrough. Variant: std (preserves MTP layer 43; one token per forward until a JANG runtime ships the accept/reject speculative decode loop). The companion DeepSeek V4 Flash JANGTQ K variant drops MTP for a smaller bundle. Routed expert layout: pre stacked along axis 0 under ffn.experts.switch mlp.{{gate proj, up proj, down proj}} per the JANGTQ PRESTACK STANDARD. Sidecar jangtq runtime.safetensors (~24 KB) ships both (in=2048, bits=2) and (in=4096, bits=2) codebooks + sign flip vectors for Swift runtimes. Bundle size: ~79.6 GB on disk Runs on: M4 Max 128 GB / M5 Max 128…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy