GME: General Multimodal Embedding GME Qwen2 VL 2B We are excited to present GME Qwen2VL series of unified multimodal embedding models , which are based on the advanced Qwen2 VL multimodal large language models (MLLMs). The GME models support three types of input: text , image , and image text pair , all of which can produce universal vector representations and have powerful retrieval performance. Key Enhancements of GME Models : Unified Multimodal Representation : GME models can process both single modal and combined modal inputs, resulting in a unified vector representation. This enables versatile retrieval scenarios (Any2Any Search), supporting tasks such as text retrieval, image retrieval from text, and image to image searches. High Performance : Achieves state of the art (SOTA) results in our universal multimodal retrieval benchmark ( UMRB ) and demonstrate strong evaluation scores in the Multimodal Textual Evaluation Benchmark ( MTEB ). Dynamic Image Resolution : Benefiting from Qwen2 VL and our training data, GME models support dynamic resolution image input. Strong Visual Retrieval Performance : Enhanced by the Qwen2 VL model series, our models excel in visual document retri…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy