This repository hosts the GGUF (llama.cpp) quantized version of MiniCPM V 4.6 Thinking. For the original BF16 weights and the full model card, please refer to openbmb/MiniCPM V 4.6 Thinking. A Pocket Sized MLLM for Ultra Efficient Image and Video Understanding on Your Phone GitHub CookBook Demo Feishu (Lark) News [2026.05.17] ⭐️⭐️⭐️ We release the API service of MiniCPM V 4.6, with a public free API key together! Try it now. MiniCPM V 4.6 Thinking MiniCPM V 4.6 Thinking is the long chain of thought reasoning variant of MiniCPM V 4.6. It generates an explicit reasoning trace before producing the final answer, substantially boosting performance on complex multimodal reasoning, math, and OCR heavy tasks, while keeping the same edge friendly architecture (SigLIP2 400M vision encoder + Qwen3.5 0.8B LLM) and the mixed 4x/16x visual token compression of MiniCPM V 4.6. Evaluation Overall Performance (Thinking) Click to view MiniCPM V 4.6 (Instruct) performance. Click to view MiniCPM V 4.6 inference efficiency results. High Concurrency Throughput Single Request TTFT (ms) Examples Overall MiniCPM V 4.6 can be deployed across three mainstream end side platforms — iOS, Android and HarmonyOS .…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy