VLM2Vec This repo contains the code and data for VLM2Vec: Training Vision Language Models for Massive Multimodal Embedding Tasks. In this paper, we aimed at building a unified multimodal embedding model for any tasks. Our model is based on converting an existing well trained VLM (Phi 3.5 V) into an embedding model. We’ve released several VLM2Vec models built on different VLM backbones: https://huggingface.co/collections/TIGER Lab/vlm2vec 6705f418271d085836e0cdd5 Also, the performance of these models is updated in the README of our GitHub repository: https://github.com/TIGER AI Lab/VLM2Vec/blob/main/README.md Release Our model is being trained on MMEB train and evaluated on MMEB eval with contrastive learning. We only use in batch negatives for training. Our best results were based on Lora training with batch size of 1024. We also have checkpoint with full training with batch size of 2048. Our results on 36 evaluation datasets are: Train/Eval Data Train data: https://huggingface.co/datasets/TIGER Lab/MMEB train Eval data: https://huggingface.co/datasets/TIGER Lab/MMEB eval VLM2Vec Checkpoints MMEB.lora8.bs1024 MMEB.fullmodel.bs2048 Github Github Experimental Results Our model can ou…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy