Model Overview Description: llama nemotron embed vl 1b v2 was developed by NVIDIA for multimodal question answering retrieval. The model can embed document pages in the form of image, text, or combined image–text inputs. Documents can be retrieved given a user query in text form. The model supports page images containing text, tables, charts, and infographics. We report the evaluation of this model on two internal multimodal retrieval benchmarks, and on the popular ViDoRe V1 and V2 benchmarks and the new Vidore V3 benchmark. An embedding model is a crucial component of a retrieval system because it transforms information into dense vector representations. An embedding model is typically a transformer encoder that processes tokens of input text or images (for example, questions, passages, or page images) to output an embedding. llama nemotron embed vl 1b v2 is a combined language and vision model. The llama nemotron embed vl 1b v2 is part of the Nemotron RAG collection of open models available on HuggingFace. It is also available for optimized inference as a NIM (NVIDIA Inference Microservice) from NVIDIA NeMo Retriever, which provides state of the art, commercially ready models and…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy