Full Modality Dataset Statistics Video Statistics Total Videos : 28,472 Total Duration : 1422.33 hours Average Duration : 179.84 seconds Median Duration : 160.08 seconds Duration Range : 10.04s 1780.03s QA Statistics Total Questions : 1,444,526 Average Questions per Video : 50.7 Questions per Video Range : 14 450 Question Type Distribution OE : 1,444,526 (100.0%) Question Category Distribution temporal : 96,873 (6.7%) causal : 96,873 (6.7%) description scene : 96,873 (6.7%) description human : 96,873 (6.7%) description object : 96,873 (6.7%) binary : 96,873 (6.7%) fine grained action understanding : 96,873 (6.7%) plot understanding : 96,873 (6.7%) non existent actions : 96,873 (6.7%) time order understanding : 96,873 (6.7%) attribute change : 96,873 (6.7%) audio visual dialogue consistency : 96,873 (6.7%) audio visual subtext : 96,873 (6.7%) audio visual mood : 96,873 (6.7%) spatial reasoning : 88,304 (6.1%) Dataset Description This dataset contains multimodal video question answering pairs that require both visual and audio information to answer correctly. The questions span multiple categories including temporal reasoning, causal analysis, scene description, and more. All questio…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy