🧠 Research & Optimization: Custom APEX Quants This repository contains custom APEX (Adaptive Precision for EXpert Models) inspired quants, tuned specifically to the underlying mixture of experts architecture and focusing on intermediary size tiers. ⚠️ Disclaimer: The layouts hosted here are custom crafted and experimental. If you are looking for the standard, stable, and predictable APEX suite of this model, please visit mudler 's excellent repository here: (https://huggingface.co/mudler/gemma 4 26B A4B it heretic APEX GGUF), where a complete baseline, uniform APEX suite is maintained. All credit for the underlying architecture, abliteration mechanics, and original model weights belongs entirely to the various upstream authors, and the imatrix dataset was generated by mradermacher. Why APEX? Regular uniform quantization applies compression equally across all layers, which can degrade MoE performance. The APEX layout locks the routing blocks at high precision and shields vital shared experts, shifting aggressive compression strictly to redundant mid layer tensors. Why custom? The standard APEX configurations generally lack an intermediary tier, between high compression Mini & Compa…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy