Qwen3.6 35B A3B PRISM NVFP4 NVFP4 (W4A4) quantization of a PRISM tuned Qwen3.6 35B A3B. ~24 GB on disk, multimodal + MTP draft head preserved. Designed for NVIDIA Blackwell (SM120/SM121). PRISM softens over refusal behaviour and removes bias / propaganda patterns while maintaining and enhancing task performance, coherence, and multimodal capability. Model details Base: Qwen/Qwen3.6 35B A3B (35B total, ~3B active per token, 256 routed experts) PRISM: refusal softening, bias + propaganda removal Format: compressed tensors NVFP4 (FP4 E2M1 weights + activations, UE4M3 per block 16 scales) Kept BF16: vision encoder, lm head , router gates, embeddings, linear attention SSM state Runtime targets: vLLM ( quantization compressed tensors ), Blackwell tensor cores Files File Purpose model.safetensors language model + vision encoder weights model mtp.safetensors MTP draft head (optional, for speculative decode) model.safetensors.index.json weight map config.json , generation config.json model + generation config tokenizer , processor config.json , chat template.jinja tokenizer + chat template Serving (vLLM) Requires vLLM with Blackwell NVFP4 kernels. On SM121 (DGX Spark), use a vLLM build with…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy