Hy3 GGUF imatrix GGUF quantizations of tencent/Hy3 — 295B total / 21B active MoE (192 experts, top 8), 80 layers + 1 MTP/NextN layer (3.8B), 256K context. Quantized day zero from the BF16 release and smoke tested on real hardware before upload (NVIDIA DGX Spark, GB10). Performance numbers below are measured, not estimated. ⚠️ Requires llama.cpp PR 25364 (unmerged) The hy v3 architecture is not yet in llama.cpp master. Until PR 25364 merges, build from the PR branch: These quants were produced at PR head a4da4b5cfdc4e5fa9def068e216a6e5154f22848 . Quants All quants use an importance matrix ( Hy3.imatrix , included) computed on a ~63KB diverse coding/reasoning/chat calibration corpus. Files are sharded at ~48GB for HF's 50GB limit — download the whole folder and point llama.cpp at the 00001 of shard; the rest load automatically. Quant Size ~BPW Fits Notes Q8 0 318 GB 8.6 server class near lossless Q5 K M 212 GB 5.8 2× 128GB class Q4 K M 181 GB 4.9 2× 128GB class recommended dual node; verified over RPC IQ4 XS 159 GB 4.3 2× 128GB class Q3 K M 143 GB 3.9 2× 128GB class IQ3 XXS 117 GB 3.2 single 128GB class (borderline) MTP block @ q8 0; best quality per GB single box Q2 K 109 GB 3.0 sin…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy