Update 2025 05 06: Replaced chat template in tokenizer config.json with the fixed version from froggeric/Qwen Fixed Chat Templates. Huihui Qwen3.6 35B A3B Claude 4.7 Opus abliterated FP8 Vision capable FP8 quantized abliterated Qwen3.6 35B A3B (MoE, hybrid mamba/attention), built from the Huihui Claude 4.7 Opus abliterated variant. Targets Nvidia DGX Spark and other FP8 capable hardware (~80 GB VRAM for full 262k context). Model Lineage Base: Qwen/Qwen3.6 35B A3B (BF16, hybrid linear attention + full attention MoE with 40 layers, 256 experts, vision encoder) Abliterated (refusals removed) by Huihui: huihui ai/Huihui Qwen3.6 35B A3B abliterated Claude 4.7 Opus style tuning by Huihui: huihui ai/Huihui Qwen3.6 35B A3B Claude 4.7 Opus abliterated This repo: FP8 quantization matching the format of Qwen/Qwen3.6 35B A3B FP8 Why FP8 Qwen3.6 35B A3B in BF16 is ~72 GB on disk. FP8 cuts that to ~37 GB while preserving vision layers and precision sensitive modules in BF16. The expected throughput uplift on DGX Spark is on par with what we saw for Qwen3.5 (31 → 51 t/s, ~65%). Quantization Details Scheme : native FP8 blockwise, identical on disk format to the official Qwen/Qwen3.6 35B A3B FP8 .…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy