OpenVid 1M — WebDataset repackaging This repository is a sequential read optimized WebDataset repackaging of nkp37/OpenVid 1M by Nan et al. (ICLR 2025). The video content is identical to the original — only the on disk layout is changed so it can be streamed efficiently from a single HTTP/NFS connection. What differs from the original Aspect Original nkp37/OpenVid 1M This repository Format Per video mp4 files zipped WebDataset .tar shards (~2 GB each) Access pattern Random per file open Sequential tar stream HF loader Custom unpacking Native load dataset(..., streaming=True) Shuffling At dataloader time Write time global shuffle + streaming buffer shuffle Metadata Separate OpenVid 1M.csv JSON sidecar per sample (same columns, spaces preserved) Integrity — manifest.json with per shard SHA 256 Re encoding or frame pre extraction were not performed — the original mp4 bytes are carried through unchanged. Statistics Train : 3,484 shards × ~2 GB ≈ 1,018,957 samples Val : 4 shards ≈ 1,000 samples (held out from the shuffled pool, fixed seed 42) Total : 1,019,957 samples Sample schema Each sample is a (mp4, json) pair inside a tar file, sharing a 9 digit key: The JSON sidecar preserves eve…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy