EoMT DINOv3 (Small, 640px) for COCO Panoptic Segmentation Overview This is the small variant of the EoMT DINOv3 model trained for panoptic segmentation on COCO at 640×640 resolution. EoMT (Encoder only Mask Transformer) is a Vision Transformer (ViT) architecture designed for high quality and efficient image segmentation. It was introduced in the CVPR 2025 highlight paper: Your ViT is Secretly an Image Segmentation Model Key Insight : Given sufficient scale and pretraining, a plain ViT along with a few additional parameters can perform segmentation without the need for task specific decoders or pixel fusion modules. The same model backbone supports semantic, instance, and panoptic segmentation with different post processing. The DINOv3 variants leverage rotary position embeddings and the latest pre training recipes from Meta AI, yielding measurable performance gains across segmentation tasks. Usage Model Details Property Value Backbone DINOv3 ViT S/16 Input Resolution 640×640 Task Panoptic Segmentation Dataset COCO Citation Acknowledgements Original implementation: tue mps/eomt Paper: arXiv:2503.19108
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy