V JEPA 2 A frontier video understanding model developed by FAIR, Meta, which extends the pretraining objectives of VJEPA, resulting in state of the art video understanding capabilities, leveraging data and model sizes at scale. The code is released in this repository. 💡 This is V JEPA 2 ViT g 384 model with video classification head pretrained on Something Something V2 dataset. Installation To run V JEPA 2 model, ensure you have installed the latest transformers: Video classification code snippet Output: Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy