Hy3 295B A21B — GGUF (IQ3 XXS UD, imatrix) GGUF quantization of tencent/Hy3 (295B total / 21B active MoE, Apache 2.0), converted from the official FP8 checkpoint (tencent/Hy3 FP8). Works with mainline llama.cpp. The hy v3 architecture was merged upstream on 2026 07 14 (PR 25395) — any master build from that date (≥ b10005) loads these files directly, including spec type draft mtp speculative decoding via the bundled NextN/MTP layer. Files file size recipe Hy3 UD128 .gguf 116.7 GB total (sharded <45 GB) routed experts: down IQ3 S · gate IQ2 S · up IQ3 XXS — attention Q5 K · shared expert + dense FFN Q6 K · output Q6 K · embeddings Q4 K Hy3 IQ4 UD split .gguf 143.9 GB total (sharded <45 GB) max quality edition for dual device rigs: routed down IQ4 XS · gate/up IQ3 S · attention/actives Q8 0 Hy3.imatrix.gguf small importance matrix, ~125 chunks of 512 tokens, general purpose calibration corpus The UD128 mix is an "unsloth dynamic" style asymmetric allocation. Tensors active on every token (attention, shared expert, dense layer 0 FFN, output head) keep high precision. Within the 192 routed experts — only 8 active per token, so they carry the low bit budget — the extra bits go to the do…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy