Llamacpp imatrix Quantizations of Mistral Small 3.2 24B Instruct 2506 by mistralai Using llama.cpp release b5697 for quantization. Original model: https://huggingface.co/mistralai/Mistral Small 3.2 24B Instruct 2506 All quants made using imatrix option with dataset from here Run them in LM Studio Run them directly with llama.cpp, or any other llama.cpp based project Prompt format What's new: Fix chat template to support tool calling Will require use of chat template and the Mistral Small 3.2 24B Instruct 2506.jinja, uploaded here and available in llama.cpp (if the PR is merged: https://github.com/ggml org/llama.cpp/pull/14349) Full server run command: Download a file (not the whole branch) from below: Filename Quant type File Size Split Description Mistral Small 3.2 24B Instruct 2506 bf16.gguf bf16 47.15GB false Full BF16 weights. Mistral Small 3.2 24B Instruct 2506 Q8 0.gguf Q8 0 25.05GB false Extremely high quality, generally unneeded but max available quant. Mistral Small 3.2 24B Instruct 2506 Q6 K L.gguf Q6 K L 19.67GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . Mistral Small 3.2 24B Instruct 2506 Q6 K.gguf Q6 K 19.35GB false Ver…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy