VideoMAE v2 (base sized model, Pretrained on UnlabeledHybrid 1M) VideoMAEv2 Base model pre trained for 800 epochs in a self supervised way on UnlabeldHybrid 1M dataset. It was introduced in the paper [[CVPR23]VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking](https://arxiv.org/abs/2203.12602) by Wang et al. and first released in GitHub. Intended uses & limitations You can use the raw model for video feature extraction. How to use Here is how to use this model to extract a video feature: BibTeX entry and citation info
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy