V JEPA 2 A frontier video understanding model developed by FAIR, Meta, which extends the pretraining objectives of VJEPA, resulting in state of the art video understanding capabilities, leveraging data and model sizes at scale. The code is released in this repository. Installation To run V JEPA 2 model, ensure you have installed the latest transformers: Intended Uses V JEPA 2 is intended to represent any video (and image) to perform video classification, retrieval, or as a video encoder for VLMs. To load a video, sample the number of frames according to the model. For this model, we use 64. To load an image, simply copy the image to the desired number of frames. For more code examples, please refer to the V JEPA 2 documentation. Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy