Kwai Keye VL 1.5 Keye VL 1.5 is a cutting edge Multimodal Large Language Model (MLLM) that addresses fundamental challenges in video comprehension. It features a novel Slow Fast video encoding strategy, a progressive four stage pre training methodology to extend context length up to 128K tokens, and a comprehensive post training pipeline focusing on reasoning enhancement and human preference alignment. The model demonstrates significant improvements in video understanding tasks and maintains competitive performance on general multimodal benchmarks. [𝕏 X] [💬 Discord] [🍎 Home Page] [📖 Technical Report] [💻 GitHub Repository] [📊 Keye VL 8B Preview ] [📊 Keye VL 1.5 8B ] [🚀 Demo] 🔥 News 2025.08.28 🌟 We are excited to introduce Kwai Keye VL 1.5 , a more powerful version! By incorporating innovative Slow Fast Video Encoding strategy , new LongCoT Cold Start data pipeline , and advanced RL training strategies , Keye VL 1.5 reaches new heights in video understanding, image comprehension, and reasoning capabilities. Plus, it now supports an extended context length of up to 128k tokens for handling longer conversations and complex tasks. Stay tuned for more groundbreaking innovations…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy