Ops MM embedding v1 7B Ops MM embedding v1 7B is a dense, large scale multimodal embedding model developed and open sourced by the Alibaba Cloud OpenSearch AI team, fine tuned from Qwen2 VL. Key Features Unified Multimodal Embeddings Encodes text, images, text image pairs, visual documents, and videos (by treating video frames as multiple image inputs) into a unified embedding space for cross modal retrieval. High Performance on MMEB Achieves SOTA results among models of similar scale on MMEB V2 and MMEB Image benchmark (until 2025 07 03). Multilingual Capabilities Ops MM embedding v1 7B achieves SOTA performance among dense models on the ViDoRe v2 benchmark, demonstrating strong cross lingual generalization. Training data MMEB train, CC 3M, colpali training set. Performance MMEB V2 Model Model Size (B) Overall Image Overall Video Overall Visdoc Overall seed 1.6 embedding unknown 71.27 77.78 55.34 73.44 Ops MM embedding v1 7B 8.29 67.61 72.72 53.76 70.34 Ops MM embedding v1 2B 2.21 63.44 69.03 47.56 66.96 VLM2Vec V2.0 Qwen2VL 2B 2.21 58.02 64.85 34.85 65.36 gme Qwen2 VL 7B Instruct 8.29 57.83 55.95 38.43 75.18 gme Qwen2 VL 2B Instruct 2.21 54.08 51.89 33.64 72.71 MMEB Image The tab…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy