A GPT 4o Level MLLM for Single Image, Multi Image and Video Understanding on Your Phone GitHub Demo MiniCPM V 4.5 MiniCPM V 4.5 is the latest and most capable model in the MiniCPM V series. The model is built on Qwen3 8B and SigLIP2 400M with a total of 8B parameters. It exhibits a significant performance improvement over previous MiniCPM V and MiniCPM o models, and introduces new useful features. Notable features of MiniCPM V 4.5 include: 🔥 State of the art Vision Language Capability. MiniCPM V 4.5 achieves an average score of 77.2 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 8B parameters, it surpasses widely used proprietary models like GPT 4o latest, Gemini 2.0 Pro, and strong open source models like Qwen2.5 VL 72B for vision language capabilities, making it the most performant MLLM under 30B parameters. 🎬 Efficient High Refresh Rate and Long Video Understanding. Powered by a new unified 3D Resampler over images and videos, MiniCPM V 4.5 can now achieve 96x compression rate for video tokens, where 6 448x448 video frames can be jointly compressed into 64 video tokens (normally 1,536 tokens for most MLLMs). This means that the model can percieve…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy