Qwen3.6 35B A3B NVFP4 NVFP4 quantized version of Qwen/Qwen3.6 35B A3B — the latest Qwen MoE with 256 experts, 3B active parameters, and state of the art coding/agentic performance. 67 GB → 21.9 GB . Single NVIDIA Blackwell GPU. 168 tok/s. Why This Model Qwen3.6 35B A3B is the new king of the MoE class: SWE bench Verified: 73.4 — surpasses models 10x its active parameter count Terminal Bench 2.0: 51.5 — best in class agentic coding QwenWebBench: 1397 ELO — real world web task performance 256 experts, 3B active — extreme sparsity = extreme speed 262K 1M context — native 262K, extensible to 1 million tokens Gated DeltaNet + Attention hybrid — next gen architecture At NVFP4, it runs at 168 tok/s on a single Blackwell GPU — faster than Gemma4 MoE (130 tok/s) with dramatically better benchmark scores. Key Specs Base model Qwen/Qwen3.6 35B A3B Architecture Qwen3.5 MoE — 35B total, 3B active , 256 experts (8 routed + 1 shared) Quantization NVFP4 W4A4 (weights FP4, activations FP4, scales FP8) Format compressed tensors (native vLLM support) Tool vllm project/llm compressor (main) Calibration 512 samples, ultrachat 200k, seq len=2048, moe calibrate all experts=True Size 21.9 GB Max context 2…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy