Llamacpp imatrix Quantizations of Phi 3.5 mini instruct Using llama.cpp release b3751 for quantization. Original model: https://huggingface.co/microsoft/Phi 3.5 mini instruct All quants made using imatrix option with dataset from here Run them in LM Studio Prompt format Download a file (not the whole branch) from below: Filename Quant type File Size Split Description Phi 3.5 mini instruct f32.gguf f32 15.29GB false Full F32 weights. Phi 3.5 mini instruct f32.gguf f32 15.29GB false Full F32 weights. Phi 3.5 mini instruct Q8 0.gguf Q8 0 4.06GB false Extremely high quality, generally unneeded but max available quant. Phi 3.5 mini instruct Q6 K L.gguf Q6 K L 3.18GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . Phi 3.5 mini instruct Q6 K.gguf Q6 K 3.14GB false Very high quality, near perfect, recommended . Phi 3.5 mini instruct Q5 K L.gguf Q5 K L 2.88GB false Uses Q8 0 for embed and output weights. High quality, recommended . Phi 3.5 mini instruct Q5 K M.gguf Q5 K M 2.82GB false High quality, recommended . Phi 3.5 mini instruct Q5 K S.gguf Q5 K S 2.64GB false High quality, recommended . Phi 3.5 mini instruct Q4 K L.gguf Q4 K L 2.47GB false…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy