[!Warning] This model has a new version: Kimi VL A3B Thinking 2506. Please consider using the new 2506 version for better abilties on general visual understanding, reasoning, video and agent scenarios. Please set a higher temperature for thinking model, especially when your problem requires relatively long thinking process. We have updated the following sample script for HF inference. 📄 Tech Report 📄 Github 💬 Chat Web 1. Introduction We present Kimi VL , an efficient open source Mixture of Experts (MoE) vision language model (VLM) that offers advanced multimodal reasoning, long context understanding, and strong agent capabilities —all while activating only 2.8B parameters in its language decoder (Kimi VL A3B). Kimi VL demonstrates strong performance across challenging domains: as a general purpose VLM, Kimi VL excels in multi turn agent interaction tasks (e.g.,OSWorld), achieving state of the art results comparable to flagship models. Furthermore, it exhibits remarkable capabilities across diverse challenging vision language tasks, including college level image and video comprehension, optical character recognition (OCR), mathematical reasoning, multi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy