nomic embed vision v1: Expanding the Latent Space nomic embed vision v1 is a high performing vision embedding model that shares the same embedding space as nomic embed text v1. All Nomic Embed Text models are now multimodal ! Name Imagenet 0 shot Datacomp (Avg. 38) MTEB : : : : : : nomic embed vision v1.5 71.0 56.8 62.28 nomic embed vision v1 70.7 56.7 62.39 OpenAI CLIP ViT B/16 68.3 56.3 43.82 Jina CLIP v1 59.1 52.2 60.1 Hosted Inference API The easiest way to get started with Nomic Embed is through the Nomic Embedding API. Generating embeddings with the nomic Python client is as easy as For more information, see the API reference Data Visualization Click the Nomic Atlas map below to visualize a 100,000 sample CC3M comparing the Vision and Text Embedding Space! Training Details We align our vision embedder to the text embedding by employing a technique similar to LiT but instead lock the text embedder! For more details, see the Nomic Embed Vision Technical Report (soon to be released!) and corresponding blog post Training code is released in the contrastors repository Usage Remember nomic embed text requires prefixes and so, when using Nomic Embed in multimodal RAG scenarios (e.g.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy