Llamacpp imatrix Quantizations of Qwen3 2507 4B Instruct Haiku 4.5 Merged by Tralalabs Using llama.cpp release b9061 for quantization. Original model: https://huggingface.co/Tralalabs/Qwen3 2507 4B Instruct Haiku 4.5 Merged FP16 All quants made using imatrix option with dataset from here Run them in your choice of tools: llama.cpp ramalama LM Studio koboldcpp Jan AI Text Generation Web UI LoLLMs Note: if it's a newly supported model, you may need to wait for an update from the developers. Prompt format Download a file (not the whole branch) from below: Filename Quant type File Size Split Description Qwen3 2507 4B Instruct Haiku 4.5 Merged bf16.gguf bf16 8.05GB false Full BF16 weights. Qwen3 2507 4B Instruct Haiku 4.5 Merged Q8 0.gguf Q8 0 4.28GB false Extremely high quality, generally unneeded but max available quant. Qwen3 2507 4B Instruct Haiku 4.5 Merged Q6 K L.gguf Q6 K L 3.57GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . Qwen3 2507 4B Instruct Haiku 4.5 Merged Q6 K.gguf Q6 K 3.48GB false Very high quality, near perfect, recommended . Qwen3 2507 4B Instruct Haiku 4.5 Merged Q5 K L.gguf Q5 K L 3.07GB false Uses Q8 0 for embed and…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy