CLIP FlanT5 XXL (VQAScore) This model is a fine tuned version of google/flan t5 xxl designed for image text retrieval tasks, as presented in the VQAScore paper. Model Description Developed by: Zhiqiu Lin and collaborators Model type: Vision Language Generative Model License: Apache 2.0 Finetuned from model: google/flan t5 xxl Model Sources [optional] Repository: https://github.com/linzhiqiu/CLIP FlanT5 Paper: https://arxiv.org/pdf/2404.01291 Demo: https://huggingface.co/spaces/zhiqiulin/VQAScore
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy