ShotBench: Expert Level Cinematic Understanding in Vision Language Models This repository contains ShotVL 3B , a fine tuned version of Qwen/Qwen2.5 VL 3B Instruct, developed for expert level cinematic understanding. Paper: ShotBench: Expert Level Cinematic Understanding in Vision Language Models Project Page: https://vchitect.github.io/ShotBench project/ Code: https://github.com/Vchitect/ShotBench Abstract Cinematography, the fundamental visual language of film, is essential for conveying narrative, emotion, and aesthetic quality. While recent Vision Language Models (VLMs) demonstrate strong general visual understanding, their proficiency in comprehending the nuanced cinematic grammar embedded within individual shots remains largely unexplored and lacks robust evaluation. This critical gap limits both fine grained visual comprehension and the precision of AI assisted video generation. To address this, we introduce ShotBench, a comprehensive benchmark specifically designed for cinematic language understanding. It features over 3.5k expert annotated QA pairs from images and video clips, meticulously curated from over 200 acclaimed (predominantly Oscar nominated) films and spanning ei…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy