Llamacpp imatrix Quantizations of gpt oss 20b heretic by p e w Using llama.cpp release b7049 for quantization. Original model: https://huggingface.co/p e w/gpt oss 20b heretic All quants made using imatrix option with dataset from here combined with a subset of combined all small.parquet from Ed Addario here Run them in LM Studio Run them directly with llama.cpp, or any other llama.cpp based project Prompt format No chat template specified so default is used. This may be incorrect, check original model card for details. Download a file (not the whole branch) from below: Filename Quant type File Size Split Description gpt oss 20b heretic bf16.gguf bf16 41.86GB false Full BF16 weights. gpt oss 20b heretic Q6 K L.gguf Q6 K L 22.19GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . gpt oss 20b heretic Q6 K.gguf Q6 K 22.19GB false Very high quality, near perfect, recommended . gpt oss 20b heretic Q5 K L.gguf Q5 K L 17.09GB false Uses Q8 0 for embed and output weights. High quality, recommended . gpt oss 20b heretic Q5 K M.gguf Q5 K M 16.90GB false High quality, recommended . gpt oss 20b heretic Q4 K L.gguf Q4 K L 16.07GB false Uses Q8 0 for em…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy