MiniCPM5 1B (GGUF Quantizations) This repository contains custom GGUF format quantizations of the openbmb/MiniCPM5 1B model. MiniCPM5 1B is a highly capable 1 billion parameter Transformer built for on device, local deployment, and resource constrained scenarios. It utilizes a standard LlamaForCausalLM architecture, features hybrid reasoning (built in tokens), and supports a massive 131k context window. 📦 Available Files and Quantizations These models were quantized specifically for high efficiency CPU/Edge inference using the llama.cpp framework. Filename Format Size Description : : : : minicpm5 1b Q4 K M.gguf Q4 K M 657 MB Excellent balance of performance and size. (Recommended for 4GB RAM/Mobile) minicpm5 1b Q5 K M.gguf Q5 K M 751 MB Higher accuracy, slight increase in size. minicpm5 1b Q6 K.gguf Q6 K 851 MB Near perfect fidelity to the base model. minicpm5 1b Q8 0.gguf Q8 0 1.1 GB Maximum quantized quality; fast loading. minicpm5 1b f16.gguf F16 2.1 GB Unquantized master weight container. 🚀 Quick Start with llama.cpp Because MiniCPM5 1B uses standard Llama architecture, it is fully supported by llama.cpp out of the box. No custom forks or kernels are required. 1. Interactive…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy