⚡ Qwen3.5 122B A10B Uncensored — APEX I Compact GGUF English 📖 中文文档 MoE Mixed Precision Quantization · Uncensored · 55.1 GB APEX I Compact MoE 122B Uncensored 55.1 GB Multimodal Qwen3.5 122B A10B (uncensored by HauhauCS ) quantized with APEX I Compact — a MoE aware mixed precision strategy that applies layer wise precision gradients. Edge layers get higher precision, middle layers get aggressive compression. Quantized from Q8 K P using the APEX project . 📊 Benchmark Results Measurements from APEX project on 8×RTX PRO 6000 Blackwell (768 GB VRAM). Perplexity on wikitext 2 raw (ctx 512). Accuracy via llama.cpp (400 tasks each). Profile Size PPL HellaSwag Wino MMLU ARC t/s Q8 0 (ref) 121 GB 4.819 85.5% 77.3% 44.19 57.19 85.5 APEX I Balanced 83.4 GB 4.831 85.5% 77.8% 43.86 57.86 96.7 APEX I Compact ★ 55.1 GB 4.978 84.5% 77.5% 44.06 57.86 106.3 APEX I Mini 44.9 GB 5.306 84.0% 75.3% 42.83 56.52 110.0 ★ This quantization . I Compact achieves 84.5% HellaSwag and 57.86 ARC at 55% less size than Q8 0 , fastest standard APEX profile at 106 t/s. Quantized from Q8 K P (137 GB → 55.1 GB). 🔬 Quantization Strategy APEX I Compact applies layer wise mixed precision with MoE aware tensor classific…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy