nomic embed text v1: A Reproducible Long Context (8192) Text Embedder Blog Technical Report AWS SageMaker Atlas Embedding and Unstructured Data Analytics Platform nomic embed text v1 is 8192 context length text encoder that surpasses OpenAI text embedding ada 002 and text embedding 3 small performance on short and long context tasks. Performance Benchmarks Name SeqLen MTEB LoCo Jina Long Context Open Weights Open Training Code Open Data : : : : : : : : : : : : : nomic embed text v1 8192 62.39 85.53 54.16 ✅ ✅ ✅ jina embeddings v2 base en 8192 60.39 85.45 51.90 ✅ ❌ ❌ text embedding 3 small 8191 62.26 82.40 58.20 ❌ ❌ ❌ text embedding ada 002 8191 60.99 52.7 55.25 ❌ ❌ ❌ Exciting Update! : nomic embed text v1 is now multimodal! nomic embed vision v1 is aligned to the embedding space of nomic embed text v1 , meaning any text embedding is multimodal! Usage Important : the text prompt must include a task instruction prefix , instructing the model which task is being performed. For example, if you are implementing a RAG application, you embed your documents as search document: and embed your user queries as search query: . Notice : From transformers v5.5.0 and sentence transformers v5.3.0,…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy