QVHighlights 1fps — Preprocessed Frames Preprocessed version of the QVHighlights dataset for temporal video grounding. Videos are extracted at 1fps , resized to 384×384 JPEG , ready for training without any video I/O at runtime. Contents File Description annotations train.jsonl 7445 train annotations annotations val.jsonl 1550 val annotations frames folder Train frames batch 0000–1000 frames 000000 001000.tar Train frames batch 0000–1000 frames 001000 002000.tar Train frames batch 1000–2000 frames 002000 003000.tar Train frames batch 2000–3000 frames 003000 004000.tar Train frames batch 3000–4000 frames 004000 005000.tar Train frames batch 4000–5000 frames 005000 006000.tar Train frames batch 5000–6000 frames 006000 007241.tar Train frames batch 6000–7241 frames 007242 007445.tar Remaining train frames val frames 000000 end.tar All val frames cache embeddings.tar embedings (siglips2 400m) of all the videos All tars extract to frames/{vid}/{timestamp:08.3f}.jpg . JSONL Format Each line is one annotation: json { "vid": "6IPwMWhdKpg 60.0 210.0", "query": "A person is riding a bike.", "duration": 150.0, "gt start sec": 12.0, "gt end sec": 45.0, "n frames": 150 } Source Annotations: jwn…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy