Model card for RAD DINO RAD DINO is a vision transformer model trained to encode chest X rays using the self supervised learning method DINOv2. Model description RAD DINO is described in detail in Exploring Scalable Medical Image Encoders Beyond Text Supervision (F. Pérez García, H. Sharma, S. Bond Taylor, et al., 2025). Developed by: Microsoft Health Futures Model type: Vision transformer License: MIT Finetuned from model: dinov2 base Uses RAD DINO is shared for research purposes only. It is not meant to be used for clinical practice . The model is a vision backbone that can be plugged to other models for downstream tasks. Some potential uses are: Image classification, with a classifier trained on top of the CLS token Image segmentation, with a decoder trained using the patch tokens Clustering, using the image embeddings directly Image retrieval, using nearest neighbors of the CLS token Report generation, with a language model to decode text Fine tuning RAD DINO is typically not necessary to obtain good performance in downstream tasks. Biases, risks, and limitations RAD DINO was trained with data from three countries, therefore it might be biased towards population in the training…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy