Qwen3.6 35B A3B — PrismaQuant 4.76 bpp Mixed precision quantization of Qwen/Qwen3.6 35B A3B produced by PrismaQuant — a per Linear sensitivity driven allocator that chooses each Linear module's format individually under a total bit budget. Why "every layer refracts into a different format": a naive uniform NVFP4 either leaves disk on the table (keeping everything BF16 "to be safe") or loses quality (quantizing sensitive layers to 4 bit). PrismaQuant measures the actual Fisher weighted MSE for every (Linear, format) pair and runs a multi choice knapsack under a total bit budget, so every bit lives where it buys the most likelihood. At a glance Metric BF16 source This artifact Delta : : : Size on disk 70 GB 22 GB −69 % Fraction of original weights 100 % 31 % Average bits per param 16 4.76 Multimodal (vision + text) ✓ ✓ MTP speculative decoding heads ✓ ✓ Loads in vLLM (stock compressed tensors ) ✓ ✓ Runtime backend any vLLM only Precision mix This checkpoint uses three precisions , selected per Linear by the allocator from measured sensitivity — not chosen uniformly: Format W A Use Count : NVFP4 4 bit (FP4, group size=16 with per group FP8 scale + per tensor global) 4 bit (dynamic) Bu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy