FLARE: Full Modality Long Video Audiovisual Retrieval Benchmark with User Simulated Queries 🤗 About This Benchmark This repository hosts the full release of FLARE: Full Modality Long Video Audiovisual Retrieval Benchmark with User Simulated Queries . FLARE screens 399 long form videos (10–60 min, 225.4 h total) from Video MME and segments them into 87,697 fine grained clips , each annotated with three captions — vision only, audio only, and unified audiovisual — and accompanied by 274,933 user simulated queries : 86,350 vision only queries rewritten from vision captions and validated by rank 1 retrieval against the vision gallery. 135,003 audio only queries rewritten from audio captions and validated by rank 1 retrieval against the audio gallery. 53,580 cross modal queries rewritten from unified captions, additionally filtered by a hard bimodal constraint — vision only retrieval fails, audio only retrieval fails, and only the joint vision+audio query uniquely identifies the target clip — so they isolate evidence that genuinely requires audiovisual fusion. Evaluation spans two axes — modality scope (vision, audio, vision+audio) and query regime (caption based, query based) — across…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy