Qwen3.6 35B A3B Uncensored Heretic MLX 4 bit · Apple Silicon native Text · Vision · Video · Thinking · Tool Calling Why this model? Three things set this apart from other Qwen 3.6 conversions: 1. Architecture aware uncensoring. Qwen 3.6 uses a hybrid attention design — linear (DeltaNet style) and traditional softmax blocks, mixed 3:1. Most abliteration tools treat them the same. llmfan46 applied separate parameters for each attention type using the Heretic tool, yielding one of the lowest KL divergences (0.0015) of any uncensored Qwen variant — 88% fewer refusals with negligible capability loss. 2. A fixed chat template. The official Qwen 3.6 template is broken on every C++ runtime (LM Studio, llama.cpp, MLX). Tool calls crash, the developer role throws errors, and empty thinking blocks waste your context window. This model ships with a rewritten template that fixes all five issues and adds a thinking toggle ( / ) you can drop into any message. 3. Vision, fixed and working. The source model had 333 vision tower keys with incorrect prefixes, breaking image inputs. Those were corrected before conversion, so text, image, and video inputs all work out…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy