A new checkpoint trained using Qwen/Qwen2 VL 2B Instruct with an enhanced training setup (LoRA tuning, batch size of 2048, maximum sub dataset size of 100k). This model has shown significantly improved performance on MMEB & Flickr30K compared to the previous models using Phi 3.5 and llava v1.6 mistral as backbone. This repo contains the code and data for VLM2Vec: Training Vision Language Models for Massive Multimodal Embedding Tasks. In this paper, we focus on building a unified multimodal embedding model suitable for a wide range of tasks. Our approach is based on transforming an existing, well trained Vision Language Model (VLM) into an embedding model. Github Github Data Our model is being trained on MMEB train and evaluated on MMEB eval with contrastive learning. We only use in batch negatives for training. Train data: https://huggingface.co/datasets/TIGER Lab/MMEB train Eval data: https://huggingface.co/datasets/TIGER Lab/MMEB eval Performance This model outperforms the baselines and previous version of VLM2Vec by a large margin. Model Classification VQA Retrieval Grounding IND OOD Overall Phi 3.5 V, Full model fine tuned ( crop=4) 52.8 50.3 57.8 72.3 62.8 47.4 55.9 Phi 3.5 V,…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy