Video Editing Understanding(VEU) Benchmark 🖥 Project Page Widely shared videos on the internet are often edited. Recently, although Video Large Language Models (Vid LLMs) have made great progress in general video understanding tasks, their capabilities in video editing understanding (VEU) tasks remain unexplored. To address this gap, in this paper, we introduce VEU Bench (Video Editing Understanding Benchmark), a comprehensive benchmark that categorizes video editing components across various dimensions, from intra frame features like shot size to inter shot attributes such as cut types and transitions. Unlike previous video editing understanding benchmarks that focus mainly on editing element classification, VEU Bench encompasses 19 fine grained tasks across three stages: recognition, reasoning, and judging. To enhance the annotation of VEU automatically, we built an annotation pipeline integrated with an ontology based knowledge base. Through extensive experiments with 11 state of the art Vid LLMs, our findings reveal that current Vid LLMs face significant challenges in VEU tasks, with some performing worse than random choice. To alleviate this issue, we develop Oscars★, a VEU e…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy