Hy3 (295B) GGUF for pulsar / ds4 / NeutronStar (SSD streaming, CUDA) Mixed precision GGUF of tencent/Hy3 (295B total / 21B active MoE, Apache 2.0) built for SSD streaming inference engines: routed experts live on disk and stream per token, so the model runs on GPUs that cannot hold it. Runs on pulsar (Rust + CUDA, recommended) and the NeutronStar hy3 branch (C, a CUDA port of antirez/ds4). Measured decode, greedy, warm cache: engine hardware tok/s pulsar RTX 5060 Ti 16GB + RTX 4060 Ti 16GB, Gen5 NVMe 7.2 pulsar RTX 4060 Ti 16GB, Gen4 NVMe 2.6 NeutronStar/ds4 RTX 4060 Ti 16GB, Gen4 NVMe 0.6 1.8 Per token only 8 of 192 experts per layer are read (~3GB/token at this quant); attention, shared experts, and the router stay resident. Files file provenance recommendation Hy3 ds4 IQ2XXS AttnQ8 fromBF16.gguf single quantization straight from the BF16 checkpoint use this one Hy3 ds4 IQ2XXS AttnQ8.gguf requantized from an IQ4 intermediate (see below) kept for continuity Both use the identical recipe and the same importance matrix; they differ only in what the quantizer saw as input. fromBF16 (new): tencent/Hy3 BF16 (598GB) converted to a q8 0 intermediate with the AngelSlim llama.cpp patches (…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy