⚡ Each donation = another big MoE quantized I host 30+ free APEX MoE quantizations as independent research. My only local hardware is an NVIDIA DGX Spark (122 GB unified memory) — enough for ~30 50B class MoEs, but bigger ones (200B+) require rented compute on H100/H200/Blackwell, typically $20 100 per quant. If APEX quants are useful to you, your support directly funds those bigger runs. 🎉 Patreon (Monthly) ☕ Buy Me a Coffee ⭐ GitHub Sponsors Qwen3.6 35B A3B Claude 4.7 Opus Reasoning Distilled — APEX MTP GGUF APEX (Adaptive Precision for EXpert Models) quantizations of lordx64/Qwen3.6 35B A3B Claude 4.7 Opus Reasoning Distilled, with the MTP (multi token prediction) head bundled for in the box self speculative decoding. Brought to you by the LocalAI team APEX Project Technical Report What's different from the plain APEX repo? These GGUFs bundle the model's MTP (multi token prediction) head alongside the trunk in a single file, courtesy of llama.cpp PR 22673. With a recent llama.cpp ( = commit 255582687) you can enable self speculative decoding using just this one file — no separate draft model needed: The non MTP version is still available at mudler/Qw…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy