Llamacpp imatrix Quantizations of Meta Llama 3.1 8B Instruct Using llama.cpp release b3472 for quantization. Original model: https://huggingface.co/meta llama/Meta Llama 3.1 8B Instruct All quants made using imatrix option with dataset from here Run them in LM Studio Torrent files https://aitorrent.zerroug.de/bartowski meta llama 3 1 8b instruct gguf torrent/ Prompt format What's new 30 07 2024: Updated chat template to fix small bug with tool usage being undefined, if you don't use the built in chat template it shouldn't change anything Download a file (not the whole branch) from below: Filename Quant type File Size Split Description Meta Llama 3.1 8B Instruct f32.gguf f32 32.13GB false Full F32 weights. Meta Llama 3.1 8B Instruct Q8 0.gguf Q8 0 8.54GB false Extremely high quality, generally unneeded but max available quant. Meta Llama 3.1 8B Instruct Q6 K L.gguf Q6 K L 6.85GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . Meta Llama 3.1 8B Instruct Q6 K.gguf Q6 K 6.60GB false Very high quality, near perfect, recommended . Meta Llama 3.1 8B Instruct Q5 K L.gguf Q5 K L 6.06GB false Uses Q8 0 for embed and output weights. High quality, r…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy