Llamacpp imatrix Quantizations of Phi 4 mini instruct by microsoft Using llama.cpp release b4792 for quantization. Original model: https://huggingface.co/microsoft/Phi 4 mini instruct All quants made using imatrix option with dataset from here Run them in LM Studio Run them directly with llama.cpp, or any other llama.cpp based project Prompt format Download a file (not the whole branch) from below: Filename Quant type File Size Split Description Phi 4 mini instruct Q8 0.gguf Q8 0 4.08GB false Extremely high quality, generally unneeded but max available quant. Phi 4 mini instruct Q6 K L.gguf Q6 K L 3.30GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . Phi 4 mini instruct Q6 K.gguf Q6 K 3.16GB false Very high quality, near perfect, recommended . Phi 4 mini instruct Q5 K L.gguf Q5 K L 3.00GB false Uses Q8 0 for embed and output weights. High quality, recommended . Phi 4 mini instruct Q5 K M.gguf Q5 K M 2.85GB false High quality, recommended . Phi 4 mini instruct Q5 K S.gguf Q5 K S 2.73GB false High quality, recommended . Phi 4 mini instruct Q4 K L.gguf Q4 K L 2.64GB false Uses Q8 0 for embed and output weights. Good quality, recommended .…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy