MotionBench: Benchmarking and Improving Fine grained Video Motion Understanding for Vision Language Models [π Project Page] [π arXiv Paper] [π Dataset] [π» GitHub] [π Leaderboard] [π HF Leaderboard] MotionBench is a comprehensive evaluation benchmark designed to assess the fine grained motion comprehension of video understanding models. It evaluates models' motion level perception through six primary categories of motion oriented question types and includes data collected from diverse sources, ensuring a broad representation of real world video content. π₯ News 2025.02.27 πππ MotionBench is accepted by CVPR 2025!! 2025.01.06 πππ We released MotionBench, a new benchmark for fine grained motion comprehension! Introduction In recent years, vision language models (VLMs) have made significant advancements in video understanding. However, a crucial capability β fine grained motion comprehension β remains under explored in current benchmarks. To address this gap, we propose MotionBench, a comprehensive evaluation benchmark designed to assess the fine grained motion comprehension of video understanding models. Features 1. Core Capabilities : Six core capabilities for fine graineβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy