MQ RAVSBench MQ RAVSBench is a benchmark for mask quality auditing in referring audio visual segmentation. Each example links a video clip, audio, a referring expression, the ground truth object mask, and candidate masks with different error patterns. The benchmark is used by MQ Auditor to assess whether a candidate mask should be accepted, revised, or rejected. All paths stored in the metadata files are relative to the dataset root. Dataset Layout Directory summary for this release: Directory Contents media/ 1,840 clips, each with audio.wav and 10 extracted frames gt mask/ Ground truth segmentation masks ( perfect ) part neg masks/ Partially incorrect masks, including cutout , erode , dilate , and merge errors full neg masks/ Masks of non target objects ( full neg ) null masks/ Empty masks used during MQ Auditor training ( null ) train test meta files/ CSV and JSON metadata used by training and evaluation scripts Metadata train test meta files/metadata.csv contains the base sample metadata: Column Meaning vid Clip id. Matches a folder under media/ uid Query/object instance id split Split label from the source metadata fid Object/mask id used in gt mask/ /fid /... exp Referring exp…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy