Llamacpp imatrix Quantizations of SmolLM3 3B by HuggingFaceTB Using llama.cpp release b5856 for quantization. Original model: https://huggingface.co/HuggingFaceTB/SmolLM3 3B All quants made using imatrix option with dataset from here Run them in LM Studio Run them directly with llama.cpp, or any other llama.cpp based project Prompt format No chat template specified so default is used. This may be incorrect, check original model card for details. What's new: Fix chat template Download a file (not the whole branch) from below: Filename Quant type File Size Split Description SmolLM3 3B bf16.gguf bf16 6.16GB false Full BF16 weights. SmolLM3 3B Q8 0.gguf Q8 0 3.28GB false Extremely high quality, generally unneeded but max available quant. SmolLM3 3B Q6 K L.gguf Q6 K L 2.59GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . SmolLM3 3B Q6 K.gguf Q6 K 2.53GB false Very high quality, near perfect, recommended . SmolLM3 3B Q5 K L.gguf Q5 K L 2.28GB false Uses Q8 0 for embed and output weights. High quality, recommended . SmolLM3 3B Q5 K M.gguf Q5 K M 2.21GB false High quality, recommended . SmolLM3 3B Q5 K S.gguf Q5 K S 2.16GB false High quality, r…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy