Model Overview Description: This model performs visual feature extraction. For instance, RADIO generates image embeddings that can be used by a downstream model to classify images. License/Terms of Use [License] This model is governed by the NVIDIA Open Model License Agreement. References: AM RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into One PHI S: Distribution Balancing for Label Free Multi Teacher Distillation RADIO Amplified: Improved Baselines for Agglomerative Vision Foundation Models Model Architecture: Architecture Type: Neural Network Network Architecture: Vision Transformer Input: Input Type(s): Image Input Format(s): Red, Green, Blue (RGB) pixel values in [0, 1] range. Input Parameters: Two Dimensional (2D) Other Properties Related to Input: Image resolutions up to 2048x2028 in increments of 16 pixels Output: Output Type(s): Embeddings Output Format: Tensor Output Parameters: 2D Other Properties Related to Output: Downstream model required to leverage image features Usage: RADIO will return a tuple with two tensors. The summary is similar to the cls token in ViT and is meant to represent the general concept of the entire image. It has shape (B,C) wi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy