E5 V: Universal Embeddings with Multimodal Large Language Models E5 V is fine tuned based on lmms lab/llama3 llava next 8b. Overview We propose a framework, called E5 V, to adpat MLLMs for achieving multimodal embeddings. E5 V effectively bridges the modality gap between different types of inputs, demonstrating strong performance in multimodal embeddings even without fine tuning. We also propose a single modality training approach for E5 V, where the model is trained exclusively on text pairs, demonstrating better performance than multimodal training. More details can be found in https://github.com/kongds/E5 V Usage Using Sentence Transformers Install Sentence Transformers: The model uses a custom chat template that automatically wraps text inputs with the instruction "Summary above sentence in one word:" and image inputs with "Summary above image in one word:". Using transformers
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy