Qwen3.6 27B — PrismaSCOUT (Blackwell, NVFP4 + BF16) PrismaQuant export of Qwen/Qwen3.6 27B for vLLM compressed tensors serving on NVIDIA Blackwell. This is the 5.31 bpp PrismaSCOUT artifact — selected by end to end held out KL on the validated Pareto frontier, replacing the prior 5.5 bpp PrismaQuant artifact. Same source weights, same export time quantization tricks (HALO, GPTQ, block output match, scale sweeps). The only thing that changed was the bit allocation routine: PrismaSCOUT selects on real measured divergence from the original model, not on summed per layer cost surrogates. Smaller and better than the prior 5.5 bpp artifact Artifact Size bpp Held out KL PrismaQuant v1 (5.5 bpp) 22.67 GB 5.50 0.0475 PrismaSCOUT (this artifact) 20.17 GB 5.31 0.0151 Change −2.5 GB (−11%) −0.19 −0.0324 (−68%) Practical impact: smaller serving footprint with substantially more VRAM headroom for KV cache, longer contexts, and concurrent request batches. About KL KL divergence (Kullback Leibler divergence) measures how different two probability distributions are. In quantization, the two distributions are: what the original full precision model would predict at a given token, and what the quanti…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy