Nemotron 3 Puzzle 75B A9B — GGUF First GGUF release of NVIDIA's Nemotron 3 Puzzle 75B A9B (hybrid mamba2/attention/latent MoE, 75B total / 9B active, 262k context, MTP draft head). Converted from the official FP8 checkpoint (weight scales absorbed at conversion — no double quantization), then quantized from the Q8 0 master with an importance matrix. Files file size note Puzzle 75B A9B Q8 0 0000X of 00002.gguf 77.7 GiB (2 shards) master, near lossless — point llama.cpp at shard 00001, the rest loads automatically Puzzle 75B A9B Q4 K M 0000X of 00002.gguf 48.1 GiB (2 shards) reference k quant, fastest decode Puzzle 75B A9B NVFP4.gguf 45.0 GiB experts NVFP4, everything else Q8 0 Puzzle 75B A9B UD IQ4 XL.gguf 41.6 GiB experts IQ4 XS; attn Q8 0, ssm/shexp Q6 K, ffn latent Q8 0 puzzle imatrix.gguf 0.2 GiB reusable imatrix (calibration datav3) Requirements Not yet supported by mainline llama.cpp — needs per layer heterogeneous MoE arrays and the 2 sub block MTP head. Use the puzzle port branch until the PR is merged: [PR LINK] Measured (Strix Halo 128GB unified, Radeon 8060S, ngl 99 ; PPL = wikitext 2 test, 24 chunks) quant PPL decode t/s prefill t/s backend Q8 0 5.325 10.2 189 Vulkan Q4…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy