📄 Tech Report 📄 Github 💬 Chat Web Introduction We present Kimi VL , an efficient open source Mixture of Experts (MoE) vision language model (VLM) that offers advanced multimodal reasoning, long context understanding, and strong agent capabilities —all while activating only 2.8B parameters in its language decoder (Kimi VL A3B). Kimi VL demonstrates strong performance across challenging domains: as a general purpose VLM, Kimi VL excels in multi turn agent interaction tasks (e.g.,OSWorld), achieving state of the art results comparable to flagship models. Furthermore, it exhibits remarkable capabilities across diverse challenging vision language tasks, including college level image and video comprehension, optical character recognition (OCR), mathematical reasoning, multi image understanding, and etc. In comparative evaluations, it effectively competes with cutting edge efficient VLMs such as GPT 4o mini, Qwen2.5 VL 7B, and Gemma 3 12B IT, while surpassing GPT 4o in several specialized domains. Kimi VL also advances the pareto frontiers of multimodal models in processing long contexts and perceiving clearly: Equipped with a 128K extended context window, K…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy