Huihui Qwen3.6 35B A3B Claude 4.7 Opus abliterated NVFP4 NVFP4 quantized version of huihui ai/Huihui Qwen3.6 35B A3B Claude 4.7 Opus abliterated — Claude 4.7 Opus distilled, abliterated (uncensored) Qwen 3.6 MoE with 256 experts and 3B active parameters. 67 GB → 21.9 GB . Single NVIDIA Blackwell GPU. 182 tok/s . 256K context. VLM. Uncensored. Why This Model Claude 4.7 Opus intelligence distilled into a locally runnable MoE, with abliteration for unrestricted research use: 256 experts, 3B active — extreme sparsity = extreme speed Claude 4.7 Opus distillation — latest Opus reasoning quality 262K native context — fits on single 96 GB GPU with FP8 KV VLM — vision fully functional (BF16 precision) Abliterated — no refusals, full capability for research and local deployment Key Specs Base model huihui ai/Huihui Qwen3.6 35B A3B Claude 4.7 Opus abliterated Architecture Qwen3.5 MoE — 35B total, 3B active , 256 experts (8 routed + 1 shared) Quantization NVFP4 W4A4 (weights FP4, activations FP4, scales FP8) Format compressed tensors (native vLLM support) Tool vllm project/llm compressor (main) Calibration 512 samples, ultrachat 200k, seq len=2048, moe calibrate all experts=True Size 21.9 GB M…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy