Llamacpp imatrix Quantizations of Hermes 4 14B by NousResearch Using llama.cpp release b6317 for quantization. Original model: https://huggingface.co/NousResearch/Hermes 4 14B All quants made using imatrix option with dataset from here combined with a subset of combined all small.parquet from Ed Addario here Run them in LM Studio Run them directly with llama.cpp, or any other llama.cpp based project Prompt format No prompt format found, check original model page Download a file (not the whole branch) from below: Filename Quant type File Size Split Description Hermes 4 14B bf16.gguf bf16 29.54GB false Full BF16 weights. Hermes 4 14B Q8 0.gguf Q8 0 15.70GB false Extremely high quality, generally unneeded but max available quant. Hermes 4 14B Q6 K L.gguf Q6 K L 12.50GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . Hermes 4 14B Q6 K.gguf Q6 K 12.12GB false Very high quality, near perfect, recommended . Hermes 4 14B Q5 K L.gguf Q5 K L 10.99GB false Uses Q8 0 for embed and output weights. High quality, recommended . Hermes 4 14B Q5 K M.gguf Q5 K M 10.51GB false High quality, recommended . Hermes 4 14B Q5 K S.gguf Q5 K S 10.26GB false High qu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy