VBVR Latent Cache (832×832 × 33f, Wan2.2 TI2V 5B VAE + UMT5 XXL) Pre encoded latent cache for the Video Reason/VBVR Dataset geometric / logical reasoning video corpus, prepared for Equilibrium Matching (EqM) post training of Wan AI/Wan2.2 TI2V 5B Diffusers on AWS Trainium2. This is a working cache, not a primary dataset. It exists to skip the ~5 s/sample VAE+T5 encode cost during training. The original videos + prompts live in the upstream VBVR Dataset repo. Source → Tensor pipeline For each VBVR sample (one directory inside an upstream tar shard), we transform two raw inputs into two encoded tensors. Nothing else from the sample dir is used (we ignore first frame.png , final frame.png , metadata.json ). What's in each .pt Key Type Shape Dtype Source x0 torch.Tensor (48, 9, 52, 52) bf16 Wan2.2 VAE latent of the resized video, normalized by per channel latents mean / latents std text embeds torch.Tensor (512, 4096) bf16 UMT5 XXL last hidden state of the prompt, padded/truncated to max seq len=512 prompt str — — Original VBVR prompt text (verbatim) source shape tuple — — (1, 3, 33, 832, 832) — the video tensor that was fed to the VAE latent shape tuple — — (48, 9, 52, 52) . Mirrors x…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy