2563eb) 8A2BE2) f9ab00) 181717?logo=github) Hy3 (Hunyuan) 1M Context GGUF : first GGUF conversion of Tencent's 299B hy v3, extended to a million tokens. Needle certified, honestly labeled. Status, precisely (updated July 9, 2026) Needle certified 10/10 at every rung from 262K to 786K (q8 KV, IQ2 M; Q4 K M also 10/10 at 524K). At the full 1,048,576 window: 70 percent per needle across three published seeds (9/10, 6/10, 6/10), so 1M is functional with real edge variance, not perfect. MTP speculative decoding was challenged during upstream review and survived a three part correctness audit including a behavioral match against Tencent's official vLLM implementation (details below). As of 2026 07 14, llama.cpp PR 25395 (hy v3 + MTP) is merged into mainline, so a recent mainline build runs these files natively no fork needed. Every claim on this card links to raw evidence in this repo. Tencent's Hy3 (299B MoE, ~17B active, hy v3 architecture) quantized to GGUF and extended to a 1,048,576 token context with YaRN. As far as we can tell these are the first GGUF quants of Hy3 anywhere: at publish time (roughly 30 hours after Hy3's release) only MLX and MXFP4 quants existed, and mainline llam…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy