Llamacpp imatrix Quantizations of OLMo 2 1124 7B Instruct Using llama.cpp release b4191 for quantization. Original model: https://huggingface.co/allenai/OLMo 2 1124 7B Instruct All quants made using imatrix option with dataset from here These were made by fixing the tokenizer.json pre processor, using the one from the base model. Prompt format Download a file (not the whole branch) from below: Filename Quant type File Size Split Description OLMo 2 1124 7B Instruct f16.gguf f16 14.60GB false Full F16 weights. OLMo 2 1124 7B Instruct Q8 0.gguf Q8 0 7.76GB false Extremely high quality, generally unneeded but max available quant. OLMo 2 1124 7B Instruct Q6 K L.gguf Q6 K L 6.19GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . OLMo 2 1124 7B Instruct Q6 K.gguf Q6 K 5.99GB false Very high quality, near perfect, recommended . OLMo 2 1124 7B Instruct Q5 K L.gguf Q5 K L 5.46GB false Uses Q8 0 for embed and output weights. High quality, recommended . OLMo 2 1124 7B Instruct Q5 K M.gguf Q5 K M 5.21GB false High quality, recommended . OLMo 2 1124 7B Instruct Q5 K S.gguf Q5 K S 5.08GB false High quality, recommended . OLMo 2 1124 7B Instruct Q4 K L.g…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy