Model Overview Description The Llama Nemotron Embedding 1B model is optimized for multilingual and cross lingual text question answering retrieval with support for long documents (up to 8192 tokens) and dynamic embedding size (Matryoshka Embeddings) . This model was evaluated on 26 languages: English, Arabic, Bengali, Chinese, Czech, Danish, Dutch, Finnish, French, German, Hebrew, Hindi, Hungarian, Indonesian, Italian, Japanese, Korean, Norwegian, Persian, Polish, Portuguese, Russian, Spanish, Swedish, Thai, and Turkish. In addition to enabling multilingual and cross lingual question answering retrieval, this model reduces the data storage footprint by 35x through dynamic embedding sizing and support for longer token length, making it feasible to handle large scale datasets efficiently. An embedding model is a crucial component of a text retrieval system, as it transforms textual information into dense vector representations. They are typically transformer encoders that process tokens of input text (for example: question, passage) to output an embedding. This model is ready for commercial use. The Llama Nemotron Embedding 1B model is a part of the NVIDIA NeMo Retriever collection o…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy