Huihui Qwen3.6 35B A3B abliterated NVFP4 NVFP4 quantized version of huihui ai/Huihui Qwen3.6 35B A3B abliterated — an abliterated (uncensored) Qwen 3.6 MoE with 256 experts, 3B active parameters, and state of the art agentic coding performance. 67 GB → 21.9 GB . Single NVIDIA Blackwell GPU. 168 tok/s. Uncensored. Why This Model All the power of Qwen3.6 35B A3B with abliteration — no refusals for local agent workflows: SWE bench Verified: 73.4 — surpasses models 10x its active size Terminal Bench 2.0: 51.5 — best in class agentic coding 256 experts, 3B active — extreme sparsity = extreme speed 262K 1M context — native 262K, extensible to 1 million tokens Abliterated — no refusals, full capability for research and local deployment Multimodal — vision preserved at full BF16 precision Key Specs Base model huihui ai/Huihui Qwen3.6 35B A3B abliterated Architecture Qwen3.5 MoE — 35B total, 3B active , 256 experts (8 routed + 1 shared) Quantization NVFP4 W4A4 (weights FP4, activations FP4, scales FP8) Format compressed tensors (native vLLM support) Tool vllm project/llm compressor (main) Calibration 512 samples, ultrachat 200k, seq len=2048, moe calibrate all experts=True Size 21.9 GB Max…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy