APEX MTP Uncensored Qwen3.6 35B A3B English π δΈζζζ‘£ π‘ What is APEX? These GGUF files are quantized using APEX , a novel MoE aware mixed precision quantization technique that outperforms standard quantization methods while being significantly smaller. APEX beats Q8 0 perplexity at half the size β and even beats F16. APEX classifies every tensor by its role β routed expert, shared expert, or attention β and applies a layer wise precision gradient, giving the most sensitive edge layers higher precision and compressing the redundant middle layers more aggressively. π¦ APEX Quantization Tiers File Size Profile Best For APEX I Quality.gguf 22 GB I Quality Highest quality, best accuracy APEX I Balanced.gguf 25 GB I Balanced Best all rounder, recommended APEX I Compact.gguf 17 GB I Compact Best quality/size ratio π Why APEX? Method Size Perplexity HellaSwag Speed F16 64.6 GB 6.537 82.5% 30.4 t/s Q8 0 34.4 GB 6.533 83.0% 52.5 t/s APEX I Quality 21.3 GB 6.552 83.5% 63.1 t/s APEX I Balanced 23.6 GB 6.548 83.0% 61.4 t/s APEX I Compact 16.1 GB 6.669 81.8% 69.8 t/s APEX Mini 12.2 GB 7.088 81.0% 74.4 t/s Benchmarks on Qwen3.5 35B A3B, NVIDIA DGX Spark (GB10, 122 GB VRAM). π I Variant: Diverseβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy