Qwen3.6 35B A3B Claude 4.7 Opus Reasoning Distilled MLX oQ4 MTP MLX/oMLX 4 bit conversion of r3lax/Qwen3.6 35B A3B Claude 4.7 Opus Reasoning Distilled with Qwen MTP tensors preserved and runtime tested in oMLX. This is not a new training run or fine tune. Weights were only converted/quantized for local MLX/oMLX inference. Quick Facts Architecture: Qwen3.6 35B A3B MoE, roughly 3B active parameters per token. Quantization: oQ4 style MLX 4 bit, group size 64. MTP: preserved and verified in oMLX. Test hardware: Apple Silicon M5 Pro with 48GB unified memory. Runtime tested with oMLX native MTP enabled. MTP Verification MTP support is runtime specific. These tensors are preserved, but non oMLX runtimes may ignore them. Local Speed Smoke Test Measured on an M5 Pro Mac with 48GB unified memory using oMLX. Per prompt MTP on smoke results: These are local smoke numbers, not a universal benchmark. Prompt, cache state, batching, oMLX version, and hardware will change results. Usage Notes Place the model under your oMLX model directory and enable native MTP: Recommended oMLX settings: Attribution Credit to the upstream work: Qwen/Qwen3.6 35B A3B for the base model. lordx64/Qwen3.6 35B A3B Claud…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy