Qwen3.6 27B — PrismaQuant 5.5 bpp Mixed precision quantization of Qwen/Qwen3.6 27B produced by PrismaQuant — a per Linear sensitivity driven allocator that chooses each Linear module's format individually under a total bit budget. Same allocator + activation aware export stack as the 35B A3B sibling; sibling coupling is pre aggregated into the DP so the achieved bpp hits the target exactly (5.500 not 5.28). This checkpoint sits at the Pareto knee of the Δloss vs bpp curve — see Why 5.5 bpp below for the full sweep and selection rationale. At a glance Metric BF16 source This artifact Delta : : : Size on disk 54 GB ~19 GB −65 % Fraction of original weights 100 % 35 % Average bits per param 16 5.50 Multimodal (vision + text) ✓ ✓ MTP speculative decoding head ✓ ✓ Loads in vLLM (stock compressed tensors ) ✓ ✓ Runtime backend any vLLM only Precision mix Selected per Linear by the allocator from measured Fisher sensitivity. On this dense 27B the allocator hit the 5.5 bpp budget exactly: Format W A Use Count (after expansion) : NVFP4 4 bit (FP4, group size=16 with per group FP8 scale + per tensor global) 4 bit (dynamic) Bulk dense MLPs + medium sensitivity attention + most visual Linears 3…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy