Hy3 — NVFP4 (routed experts, MSE scales) ⚠️ Known behavior reports (under investigation): Community testing has surfaced occasional Chinese output in English contexts (reported in the full precision preview as well, so at least partly a base model trai: uncalibrated fp8 KV appears to amplify it) and intermittent tool call failures / premature stops, with chat template interaction as the current suspect. Until the KV calibrated revision lands, use bf16 KV ( kv cache dtype auto) and num speculative tokens 1 on GB10 class hardware. Differential evals against the BF16 baseline are in progress; results will be published here. A 4 bit NVFP4 quantization of tencent/Hy3 . The original model card is preserved in full below. 2x GB10 full recipe, scripts, and all the bugs documented by Tony DeAngelo (tonyd2wild): https://github.com/tonyd2wild/Hy3 295B NVFP4 MTP 2x DGX Spark Weight only NVFP4 quant produced with qstream using MSE optimal group scale selection: the routed MoE experts (≈95% of the weights) are quantized to 4 bit; everything quality sensitive stays BF16. ⚠️ MARLIN only (this is a W4A16 build) Because the activations stay BF16 ( weight only, W4A16 ), vLLM serves this build exclusi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy