⚡ Each donation = another big MoE quantized I host 25+ free APEX MoE quantizations as independent research. My only local hardware is an NVIDIA DGX Spark (122 GB unified memory) — enough for ~30 50B class MoEs, but bigger ones (200B+) require rented compute on H100/H200/Blackwell, typically $20 100 per quant. If APEX quants are useful to you, your support directly funds those bigger runs. 🎉 Patreon (Monthly) ☕ Buy Me a Coffee ⭐ GitHub Sponsors 💚 Big thanks to Hugging Face for generously donating additional storage — much appreciated. Qwen3.5 122B A10B APEX GGUF APEX (Adaptive Precision for EXpert Models) quantizations of Qwen3.5 122B A10B. Brought to you by the LocalAI team APEX Project Technical Report Benchmark Results All measurements on 8xRTX PRO 6000 Blackwell (768 GB VRAM). Perplexity on wikitext 2 raw, context 512. Accuracy benchmarks via llama.cpp (400 tasks each). Configuration Size (GB) Perplexity KL mean HellaSwag Winogrande MMLU ARC tg128 (t/s) Q8 0 (Unsloth) 121 4.819 0.004 85.5% 77.3% 44.19 57.19 85.5 Q5 K S (Unsloth) ~81 4.826 0.007 85.3% 76.0% 43.80 57.86 90.4 UD Q4 K XL (Unsloth) ~72 4.829 0.010 84.8% 76.3% 44.25 55.85 91.8 APEX I Bala…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy