FastVideo Synthetic Wan2.2 720P dataset FastVideo Team Paper Github Project Page Abstract Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware efficient sparse attention that replaces full attention at \emph{both} training and inference. In VSA, a… See the full description on the dataset page: https://huggingface.co/datasets/FastVideo/Wan2.2 Syn 121x704x1280 32k.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy