Model Card for EUPE Running AI models on smart edge devices can unlock various user experiences, but presents challenges due to limited compute and the need to handle multiple tasks simultaneously. This requires a vision encoder with small size but powerful and versatile representations. We present our method, Efficient Universal Perception Encoder (EUPE), which offers both inference efficiency and universally good representations for diverse downstream tasks. We achieve this by distilling from multiple domain expert foundation vision encoders. Unlike previous agglomerative methods that directly scale down from multiple teachers to an efficient encoder, we demonstrate the importance of first scaling up to a large proxy teacher and then distilling from this single teacher. Experiments show that EUPE achieves on par or better performance than individual domain experts of the same size on diverse task domains and also outperforms previous agglomerative encoders. Model Details These are Vision Transformer and ConvNeXt models trained following the method described in the EUPE paper. 6 models are provided: 3 ViT models including ViT B16, ViT S16, ViT T16 3 ConvNeXt models including ConvN…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy