The embedding set trained by Jina AI . Jina CLIP: your CLIP model is also your text retriever! Intended Usage & Model Info jina clip v1 is a state of the art English multimodal (text image) embedding model . Traditional text embedding models, such as jina embeddings v2 base en, excel in text to text retrieval but incapable of cross modal tasks. Models like openai/clip vit base patch32 effectively align image and text embeddings but are not optimized for text to text retrieval due to their training methodologies and context limitations. jina clip v1 bridges this gap by offering robust performance in both domains. Its text component matches the retrieval efficiency of jina embeddings v2 base en , while its overall architecture sets a new benchmark for cross modal retrieval. This dual capability makes it an excellent tool for multimodal retrieval augmented generation (MuRAG) applications, enabling seamless text to text and text to image searches within a single model. Data & Parameters Check out our paper Usage 1. The easiest way to starting using jina clip v1 en is to use Jina AI's Embeddings API. 2. Alternatively, you can use Jina CLIP directly via transformers/sentence transformers…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy