Kwai Keye VL [𝕏 X] [💬 Discord] [🍎 Home Page] [📖 Technical Report] [💻 GitHub Repository] [📊 Keye VL 8B Preview ] [📊 Keye VL 1.5 8B ] [🚀 Demo] Abstract While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information dense short form videos, a dominant medium in today's digital landscape. To bridge this gap, we introduce Kwai Keye VL , an 8 billion parameter multimodal foundation model engineered for leading edge performance in short video understanding while maintaining robust general purpose vision language abilities. The development of Keye VL rests on two core pillars: a massive, high quality dataset exceeding 600 billion tokens with a strong emphasis on video, and an innovative training recipe. This recipe features a four stage pre training process for solid vision language alignment, followed by a meticulous two phase post training process. The first post training stage enhances foundational capabilities like instruction following, while the second phase focuses on stimulating advanced reasoning. In this second phase, a key innovation is our five mode cold start'' data mixtur…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy