rinna/japanese clip vit b 16 This is a Japanese CLIP (Contrastive Language Image Pre Training) model trained by rinna Co., Ltd.. Please see japanese clip for the other available models. How to use the model 1. Install package 2. Run Model architecture The model was trained a ViT B/16 Transformer architecture as an image encoder and uses a 12 layer BERT as a text encoder. The image encoder was initialized from the AugReg vit base patch16 224 model. Training The model was trained on CC12M translated the captions to Japanese. Release date May 12, 2022 How to cite License The Apache 2.0 license
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy