The embedding set trained by Jina AI . Jina CLIP v2: Multilingual Multimodal Embeddings for Texts and Images This model is based on the paper jina clip v2: Multilingual Multimodal Embeddings for Text and Images. Quick Start Blog Technical Report Azure AWS SageMaker Google Cloud Platform API Intended Usage & Model Info jina clip v2 is a general purpose multilingual multimodal embedding model for text & images . Multimodal embeddings enable searching and understanding data across different modalities through a coherent representation. They serve as the backbone of neural information retrieval and multimodal GenAI applications. Built upon jina clip v1 and our recently released jina embeddings v3 , jina clip v2 features several significant improvements: Improved Performance : v2 shows a 3% performance improvement over v1 in both text image and text text retrieval tasks. Similar to v1, v2's text encoder can serve as an effective multilingual long context dense retriever. It performs on par with our frontier model jina embeddings v3 (currently the best multilingual embeddings under 1B parameters on MTEB). Multilingual Support : Using the same backbone as jina embeddings v3 for the text t…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy