⚡ Each donation = another big MoE quantized I host 25+ free APEX MoE quantizations as independent research. My only local hardware is an NVIDIA DGX Spark (122 GB unified memory) — enough for ~30 50B class MoEs, but bigger ones (200B+) require rented compute on H100/H200/Blackwell, typically $20 100 per quant. If APEX quants are useful to you, your support directly funds those bigger runs. 🎉 Patreon (Monthly) ☕ Buy Me a Coffee ⭐ GitHub Sponsors 💚 Big thanks to Hugging Face for generously donating additional storage — much appreciated. Qwen3.5 35B A3B APEX GGUF A Novel MoE Aware Mixed Precision Quantization Technique Brought to you by the LocalAI team the creators of LocalAI the open source AI engine that runs any model LLMs, vision, voice, image, video on any hardware. No GPU required. APEX Technical Report GitHub Repository LocalAI APEX (Adaptive Precision for EXpert Models) is a novel quantization technique for Mixture of Experts language models. Unlike uniform quantization methods that apply the same precision to every tensor, APEX introduces a layer wise precision gradient combined with MoE aware tensor classification and diverse imatrix calibration…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy