CrossVid: A Comprehensive Benchmark for Evaluating Cross Video Reasoning in Multimodal Large Language Models Dataset Description CrossVid is a large scale, multi task dataset designed to advance cross video understanding capabilities in vision language models. The dataset encompasses 10 diverse task types that require models to reason across multiple videos, understand temporal dynamics, spatial relationships, and complex narrative structures. Unlike existing benchmarks focusing on single video analysis, CrossVid is the first comprehensive benchmark designed to evaluate cross video understanding capabilities in MLLMs. Key Features 🎥 Multi Domain Videos : Includes assembly tutorials, animal/human behaviors, cooking demonstrations, movie scenes, and UAV footage 🎯 10 Challenging Tasks : Covering behavioral analysis, content comparison, temporal reasoning, spatial understanding, and more 📊 Rich Annotations : Question answer pairs with temporal segments, spatial object tracking, and procedural step sequences 🌐 Cross Video Reasoning : Tasks explicitly require understanding relationships and patterns across multiple video clips Task Types Task Code Task Name Dimension QA Pairs Videos…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy