V JEPA 2 A frontier video understanding model developed by FAIR, Meta, which extends the pretraining objectives of VJEPA, resulting in state of the art video understanding capabilities, leveraging data and model sizes at scale. The code is released in this repository. Installation To run V JEPA 2 model, ensure you have installed the latest transformers: Intended Uses V JEPA 2 is intended to represent any video (and image) to perform video classification, retrieval, or as a video encoder for VLMs. To load a video, sample the number of frames according to the model. For this model, we use 64. To load an image, simply copy the image to the desired number of frames. For more code examples, please refer to the V JEPA 2 documentation. Citation @techreport{assran2025vjepa2, title={V JEPA~2: Self Supervised Video Models Enable Understanding, Prediction and Planning}, author={Assran, Mahmoud and Bardes, Adrien and Fan, David and Garrido, Quentin and Howes, Russell and Komeili, Mojtaba and Muckley, Matthew and Rizvi, Ammar and Roberts, Claire and Sinha, Koustuv and Zholus, Artem and Arnaud, Sergio and Gejji, Abha and Martin, Ada and Robert Hogan, Francois and Dugas, Daniel and Bojanow…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy