Model Card for PubMedCLIP PubMedCLIP is a fine tuned version of CLIP for the medical domain. Model Description PubMedCLIP was trained on the Radiology Objects in COntext (ROCO) dataset, a large scale multimodal medical imaging dataset. The ROCO dataset includes diverse imaging modalities (such as X Ray, MRI, ultrasound, fluoroscopy, etc.) from various human body regions (such as head, spine, chest, abdomen, etc.) captured from open access PubMed articles. PubMedCLIP was trained for 50 epochs with a batch size of 64 using the Adam optimizer with a learning rate of 10−5. The authors have released three different pre trained models at this link which use ResNet 50, ResNet 50x4 and ViT32 as image encoders. This repository includes only the ViT32 variant of the PubMedCLIP model. Repository: PubMedCLIP Official GitHub Repository Paper: Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain? Usage Additional Information Licensing Information The authors have released the model code and pre trained checkpoints under the MIT License. Citation Information
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy