Built for vMLX — the MLX inferencer with VL + video, KV cache quantization, prefix cache reuse, agentic tool calling, and native MTP speculative decoding. Free for macOS · vmlx.net Qwen 3.6 35B A3B — MXFP4 CRACK + d3 MTP CRACK abliterated · MXFP4 (4 bit microscaling) · d3 MTP self speculative (1.51× faster) · Vision + Video · Reasoning toggle · 18 GB What Is This? This is Qwen 3.6 35B A3B — a vision language model (Mixture of Experts (256 routed, 10 active) hybrid SSM + full attention, 40 layers, native image + video understanding) that has been: 1. CRACK abliterated — refusal behavior removed at the weight level, so it complies across task categories instead of refusing, while keeping its knowledge, reasoning, and vision intact. 2. MXFP4 (4 bit microscaling) quantized for MLX on Apple Silicon — 18 GB. 3. MTP preserved — the native multi token prediction head is kept and abliterated too, so d3 self speculative decoding works (~1.51× faster) on an MTP aware runtime (vMLX). Vision and video processing are fully preserved. Results Evaluated through the vMLX inference engine. HarmBench scored with a strict classifier (rejects loops, empty/template dumps, and thinking trace leakage). MM…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy