⚡ Each donation = another big MoE quantized I host 25+ free APEX MoE quantizations as independent research. My only local hardware is an NVIDIA DGX Spark (122 GB unified memory) — enough for ~30 50B class MoEs, but bigger ones (200B+) require rented compute on H100/H200/Blackwell, typically $20 100 per quant. If APEX quants are useful to you, your support directly funds those bigger runs. 🎉 Patreon (Monthly) ☕ Buy Me a Coffee ⭐ GitHub Sponsors 💚 Big thanks to Hugging Face for generously donating additional storage — much appreciated. Nemotron 3 Nano 30B A3B APEX GGUF APEX (Adaptive Precision for EXpert Models) quantizations of NVIDIA Nemotron 3 Nano 30B A3B. Brought to you by the LocalAI team APEX Project Technical Report Benchmark Results Benchmarks coming soon. For reference APEX benchmarks on the Qwen3.5 35B A3B architecture, see mudler/Qwen3.5 35B A3B APEX GGUF. What is APEX? APEX is a quantization strategy for Mixture of Experts (MoE) models. It classifies tensors by role (routed expert, shared expert, attention) and applies a layer wise precision gradient edge layers get higher precision, middle layers get more aggressive compression. I variants…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy