nomic embed text v1: A Reproducible Long Context (8192) Text Embedder nomic embed text v1 is 8192 context length text encoder that surpasses OpenAI text embedding ada 002 and text embedding 3 small performance on short and long context tasks. Name SeqLen MTEB LoCo Jina Long Context Open Weights Open Training Code Open Data : : : : : : : : : : : : : nomic embed text v1 8192 62.39 85.53 54.16 ✅ ✅ ✅ jina embeddings v2 base en 8192 60.39 85.45 51.90 ✅ ❌ ❌ text embedding 3 small 8191 62.26 82.40 58.20 ❌ ❌ ❌ text embedding ada 002 8191 60.99 52.7 55.25 ❌ ❌ ❌ Hosted Inference API The easiest way to get started with Nomic Embed is through the Nomic Embedding API. Generating embeddings with the nomic Python client is as easy as For more information, see the API reference Data Visualization Click the Nomic Atlas map below to visualize a 5M sample of our contrastive pretraining data! Training Details We train our embedder using a multi stage training pipeline. Starting from a long context BERT model, the first unsupervised contrastive stage trains on a dataset generated from weakly related text pairs, such as question answer pairs from forums like StackExchange and Quora, title body pairs fro…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy