Qwen3.6 35B A3B — Claude 4.7 Opus Reasoning Distilled — PrismaQuant 4.75 bpp Mixed precision quantization of Qwen3.6 35B A3B Claude 4.7 Opus Reasoning Distilled produced by PrismaQuant — a per Linear sensitivity driven allocator that chooses each Linear module's format individually under a total bit budget. The source model is a reasoning distillation of Qwen3.6 35B A3B trained on Claude 4.7 Opus chain of thought traces. It retains the full multimodal (vision + text) and MTP speculative decoding head of the base architecture, with added reasoning capability. Enable thinking mode at inference time (see Serving below). Why "every layer refracts into a different format": a naive uniform NVFP4 either leaves disk on the table (keeping everything BF16 "to be safe") or loses quality (quantizing sensitive layers to 4 bit). PrismaQuant measures the actual Fisher weighted MSE for every (Linear, format) pair and runs a multi choice knapsack under a total bit budget, so every bit lives where it buys the most likelihood. At a glance Metric BF16 source This artifact Size on disk ~70 GB ~22 GB Average bits per param 16 4.75 Reasoning (Claude 4.7 Opus distillation) ✓ ✓ Multimodal (vision + text) ✓…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy