⚡ Each donation = another big MoE quantized I host 25+ free APEX MoE quantizations as independent research. My only local hardware is an NVIDIA DGX Spark (122 GB unified memory) — enough for ~30 50B class MoEs, but bigger ones (200B+) require rented compute on H100/H200/Blackwell, typically $20 100 per quant. If APEX quants are useful to you, your support directly funds those bigger runs. 🎉 Patreon (Monthly) ☕ Buy Me a Coffee ⭐ GitHub Sponsors 💚 Big thanks to Hugging Face for generously donating additional storage — much appreciated. Qwen 3.6 35B A3B APEX GGUF APEX (Adaptive Precision for EXpert Models) quantizations of Qwen/Qwen3.6 35B A3B. Brought to you by the LocalAI team APEX Project Technical Report Benchmark Results All benchmarks run with llama.cpp b8797 on NVIDIA GB10 (122 GB VRAM). Perplexity and KL divergence measured on wikitext 2. HellaSwag zero shot (400 tasks). KL divergence computed against BF16 reference logits. APEX vs Baselines (unsloth UD quants) Model Size PPL ↓ KL mean ↓ KL median ↓ KL max ↓ HellaSwag ↑ BF16 (reference) 65 GB 6.722 — — — — Q8 0 35 GB 6.720 0.0059 0.0022 9.72 82.5% UD Q5 K XL 25 GB 6.725 0.0083 0.0030 9.06 82.8% UD…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy