A Gemini 2.5 Flash Level MLLM for Vision, Speech, and Full Duplex Mulitmodal Live Streaming on Your Phone GitHub CookBook Omni modal Demo Vision Language Demo WeChat Discord CaseBook(Audio, Omni Full Duplex) News [2026.05.17] ⭐️⭐️⭐️ We release the API service of MiniCPM o 4.5, supporting both traditional text and vision language requests, and also full duplex realtime interaction! Try it now. [2026.02.06] 🥳 🥳 🥳 We open sourced a realtime web demo deployable on your own devices like Mac or GPU. Try it now! MiniCPM o 4.5 MiniCPM o 4.5 is the latest and most capable model in the MiniCPM o series. The model is built in an end to end fashion based on SigLip2, Whisper medium, CosyVoice2, and Qwen3 8B with a total of 9B parameters. It exhibits a significant performance improvement, and introduces new features for full duplex multimodal live streaming. Notable features of MiniCPM o 4.5 include: 🔥 Leading Visual Capability. MiniCPM o 4.5 achieves an average score of 77.6 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. With only 9B parameters, it surpasses widely used proprietary models like GPT 4o, Gemini 2.0 Pro, and approaches Gemini 2.5 Flash for vision language c…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy