LLaVA OneVision 2 Data Training data for the LLaVA OneVision 2 multimodal model family, covering large scale video and spatial reasoning corpora used in mid training. Dataset Composition Subset Format Description mid training video/60s rest/ WebDataset ( .tar ) 10,809 shards of ~60s video clips mid training video/caption v0/split 30s.jsonl JSONL Captions for 30 second video clips mid training video/caption v0/split 60s.jsonl JSONL Captions for 60 second video clips mid training video/caption v0/split 180s.jsonl JSONL Captions for 180 second video clips mid training video/caption v0/split gt10min.jsonl JSONL Captions for 10 minute video clips spatial/ WebDataset ( .tar ) 84 shards of spatial reasoning data (refcoco, visual genome, pointing, 3D, etc.) mid training video/mapping/mapping {5s,10s,30s,60s,180s,gt10min}.csv CSV Maps each video clip's dst path to its source youtube id and [start time, end time] window Preview Configs The viewer configs above expose small Parquet samples so the Hugging Face Dataset Viewer can render the data directly in the browser: viewer caption 30s — 5 caption samples from 30 second clips viewer caption 60s — 5 caption samples from 60 second clips viewer…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy