⚡ Each donation = another big MoE quantized I host 25+ free APEX MoE quantizations as independent research. My only local hardware is an NVIDIA DGX Spark (122 GB unified memory) — enough for ~30 50B class MoEs, but bigger ones (200B+) require rented compute on H100/H200/Blackwell, typically $20 100 per quant. If APEX quants are useful to you, your support directly funds those bigger runs. 🎉 Patreon (Monthly) ☕ Buy Me a Coffee ⭐ GitHub Sponsors 💚 Big thanks to Hugging Face for generously donating additional storage — much appreciated. Gemma 4 26B A4B APEX GGUF APEX (Adaptive Precision for EXpert Models) quantizations of google/gemma 4 26B A4B it. Brought to you by the LocalAI team APEX Project Technical Report Benchmark Results Benchmarks coming soon (re quantized with llama.cpp b8664 including Gemma 4 tokenizer and logit softcapping fixes). For reference APEX benchmarks on the Qwen3.5 35B A3B architecture, see mudler/Qwen3.5 35B A3B APEX GGUF. Available Files File Profile Size Best For gemma 4 26B A4B APEX I Balanced.gguf I Balanced 19 GB Best overall quality/size ratio gemma 4 26B A4B APEX I Quality.gguf I Quality 20 GB Highest quality with imatrix ge…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy