Llamacpp imatrix Quantizations of Llama 3 3 Nemotron Super 49B v1 5 by nvidia Using llama.cpp release b5962 for quantization. Original model: https://huggingface.co/nvidia/Llama 3 3 Nemotron Super 49B v1 5 All quants made using imatrix option with dataset from here Run them in LM Studio Run them directly with llama.cpp, or any other llama.cpp based project Prompt format Download a file (not the whole branch) from below: Filename Quant type File Size Split Description Llama 3 3 Nemotron Super 49B v1 5 bf16.gguf bf16 99.74GB true Full BF16 weights. Llama 3 3 Nemotron Super 49B v1 5 Q8 0.gguf Q8 0 52.99GB true Extremely high quality, generally unneeded but max available quant. Llama 3 3 Nemotron Super 49B v1 5 Q6 K L.gguf Q6 K L 41.43GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . Llama 3 3 Nemotron Super 49B v1 5 Q6 K.gguf Q6 K 40.92GB false Very high quality, near perfect, recommended . Llama 3 3 Nemotron Super 49B v1 5 Q5 K L.gguf Q5 K L 36.04GB false Uses Q8 0 for embed and output weights. High quality, recommended . Llama 3 3 Nemotron Super 49B v1 5 Q5 K M.gguf Q5 K M 35.39GB false High quality, recommended . Llama 3 3 Nemotron Supe…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy