Llamacpp imatrix Quantizations of gpt oss 20b by openai Using llama.cpp release b6096 for quantization. Original model: https://huggingface.co/openai/gpt oss 20b All quants made using imatrix option with combined all medium dataset from Ed Addario here Run them in LM Studio Run them directly with llama.cpp, or any other llama.cpp based project All quants keep the feed forward networks at mxfp4 for optimal performance, which does mean the size differences are negligible unfortunately, but being provided just because. Prompt format No chat template specified so default is used. This may be incorrect, check original model card for details. Download a file (not the whole branch) from below: Use this one: Filename Quant type File Size Split Description gpt oss 20b MXFP4.gguf MXFP4 12.1GB false Full MXFP4 weights, recommended for this model. The reason is, the FFN (feed forward networks) of gpt oss do not behave nicely when quantized to anything other than MXFP4, so they are kept at that level for everything. The rest of these are provided for your own interest in case you feel like experimenting, but the size savings is basically non existent so I would not recommend running them, they…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy