LLaVA-OneVision-2-Data
Training data for the LLaVA-OneVision-2 multimodal model family, covering large-scale video and spatial reasoning corpora used in mid-training.
Dataset Composition
| Subset | Format | Description |
|---|---|---|
mid_training_video/60s_rest/ | WebDataset (.tar) | 10,809 shards of ~60s video clips |
mid_training_video/caption_v0/split_30s.jsonl | JSONL | Captions for 30-second video clips |
mid_training_video/caption_v0/split_60s.jsonl | JSONL | Captions for 60-second video clips |
mid_training_video/caption_v0/split_180s.jsonl | JSONL | Captions for 180-second video clips |
mid_training_video/caption_v0/split_gt10min.jsonl | JSONL | Captions for >10-minute video clips |
spatial/ | WebDataset (.tar) | 84 shards of spatial reasoning data (refcoco, visual genome, pointing, 3D, etc.) |
mid_training_video/mapping/mapping_{5s,10s,30s,60s,180s,gt10min}.csv | CSV | Maps each video clip's dst_path to its source youtube_id and [start_time, end_time] window |
Preview Configs
The viewer_* configs above expose small Parquet samples so the Hugging Face Dataset Viewer can render the data directly in the browser:
viewer_caption_30s— 5 caption samples from 30-second clipsviewer_caption_60s— 5 caption samples from 60-second clipsviewer_caption_180s— 3 caption samples from 180-second clipsviewer_caption_gt10min— 1 caption sample from >10-minute clipsviewer_spatial— 10 spatial-reasoning samples with embedded thumbnail images, mixed across tasks (refcoco, visual genome, pointing, ca1m, osd, crosspoint, erqa, roborefer)
These previews are intended for schema inspection only. For training, use the full mid_training_video/ and spatial/ shards.