NVIDIA Nemotron Labs 3 Puzzle 75B A9B GGUF Community GGUF conversion of nvidia/NVIDIA Nemotron Labs 3 Puzzle 75B A9B BF16 . Important llama.cpp compatibility notice This model was converted with the still unmerged ggml org/llama.cpp PR 25444, pinned to commit af49ef5cd990d039dbf360dd3a9f3b5dafdd1726 , plus a narrowly scoped converter compatibility fix for the official BF16 checkpoint's model.layers. tensor prefix and bounded writeback for large lazy tensors on the high RAM Colab runtime. Until equivalent support is merged into mainline llama.cpp, use a build containing PR 25444 to load these files. PR 25444 adds NemotronHPuzzleForCausalLM / nemotron h puzzle support, heterogeneous per layer MoE settings, and the model's two block MTP draft head. This repository is not an official NVIDIA or llama.cpp release. Files NVIDIA Nemotron Labs 3 Puzzle 75B A9B BF16.gguf — BF16 master GGUF NVIDIA Nemotron Labs 3 Puzzle 75B A9B Q2 K.gguf NVIDIA Nemotron Labs 3 Puzzle 75B A9B Q3 K S.gguf NVIDIA Nemotron Labs 3 Puzzle 75B A9B Q3 K M.gguf NVIDIA Nemotron Labs 3 Puzzle 75B A9B Q3 K L.gguf NVIDIA Nemotron Labs 3 Puzzle 75B A9B IQ4 XS.gguf NVIDIA Nemotron Labs 3 Puzzle 75B A9B Q4 K S.gguf NVIDIA Ne…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy