Huihui Qwen3.6 35B A3B abliterated FP8 DYNAMIC FP8 dynamic quantization of huihui ai/Huihui Qwen3.6 35B A3B abliterated, produced for inference on NVIDIA DGX Spark (GB10, SM121) under its 128 GB unified memory budget. This is the clean abliteration variant quantized to FP8, distinct from existing Claude distilled variants on the Hub. Uses llm compressor FP8 DYNAMIC scheme + compressed tensors float quantized format. ⚠️ Vision pipeline fix (2026 05 03) Pre 2026 05 03 uploads of this checkpoint had broken vision input — the model would emit !!!!! token loops on any image. Root cause: vLLM (≤0.20.0) expects vision tensor keys at visual. , but Qwen3 5MoeForConditionalGeneration saves them under model.language model.visual. , so all 333 vision tower tensors got silently skipped during load. This is now fixed : shard 19 + model.safetensors.index.json re uploaded with the visual prefix stripped. Text only output is unaffected. If you cloned the model before 2026 05 03, the cleanest path is huggingface hub.snapshot download(..., force download=True, allow patterns=["model 00019 of 00019.safetensors", "model.safetensors.index.json"]) — only ~600 MB of changed weights. Diagnosis + remap scri…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy