VidLayer Dataset π Introduction We introduce VidLayer , a large scale dataset specifically designed for layer aware video generation. VidLayer provides aligned foreground videos, foreground masks, background videos, and raw videos, enabling supervision at both the semantic and structural levels. We believe constructing VidLayer is a foundational step toward layer wise text to video modeling. π Dataset Details We organize the VidLayer dataset according to different sources: VIDGEN GOT 10K Youtube VOS MOSEv1 MOSEv2 For each data point, we have: 1. Original Video Clip (BG + FG) 2. Background Video Clip (Pure BG) 3. Foreground Object Mask 4. (optional) Foreground Video Clip (Pure FG) 5. Multi Layer Prompt For data point that does not include foreground video clip, you can simply use FG = Video Mask to obtain the FG. Data Construction Pipeline: Data Visualization Samples: β¨ Citation If you find this dataset or the associated work useful for your research, please cite the paper:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy