ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models 📄 arXiv Paper 🖥️ Github Code 📦 Data Introduction ProactiveVideoQA is the first comprehensive benchmark designed to evaluate a system's ability to engage in proactive interaction in multimodal dialogue settings. Unlike traditional turn by turn dialogue systems, in proactive intraction model need to determine when to repsond during the playback, so both response timing and response textual content are important points for evaluation. Dataset Statistics ProactiveVideoQA contains 4 tasks: 1. Proactive web video QA [WEB] : centering on general web video understanding. 1. Proactive ego centric video QA [EGO] : centering on first person view video comprehension, particularly relevant in robotics and daily assistant applications. 1. Proactive TV series video QA [TV] : emphasizing dialogue and social relationship understanding with speech input, and 1. Proactive video anomaly detection [VAD] targeting surveillance video monitoring and alerting. 1377 videos from different sources 1427 different qeustions, and 3510 ground truth reply turns Fully proactive questions and open ended an…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy