VidEoMT S on YouTube VIS 2019 This repository contains the Hugging Face Transformers conversion of the official VidEoMT checkpoint yt 2019 vit small 52.8.pth from tue mps/VidEoMT. Model details Architecture: VidEoMT with a DINOv2 ViT S/14 with 4 register tokens backbone Task: video instance segmentation Dataset: YouTube VIS 2019 Input resolution: 640 x 640 Number of frames: 2 Paper: Your ViT is Secretly Also a Video Segmentation Model Reported metrics Metric Value AP 52.8 AR@10 62.2 FPS 294 The metrics above are the numbers reported by the authors in the official model zoo. Usage Use processor.post process instance segmentation , processor.post process panoptic segmentation , or processor.post process semantic segmentation depending on the target task.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy