Omni Embed Nemotron 3B Description NV QwenOmni Embed 3B v1 is a versatile multimodal embedding model capable of encoding content across multiple modalities, including text, image, audio, and video, either individually or in combination, and supports retrieval using queries that can also be multimodal. It is designed to serve as a foundational component in multi modal Retrieval Augmented Generation (RAG) systems. The foundational Qwen Omni model (Qwen/Qwen2.5 Omni 3B) is based on the Thinker Talker architecture. We only leverage the Thinker component to encode and understand diverse modalities. In this implementation, we do not include the Talker component, as the model focuses on multimodal understanding rather than response generation. This model is for research and development only. For more technical details, please refer to our technical report: Omni Embed Nemotron: A Unified Multimodal Retrieval Model for Text, Image, Audio, and Video License/Terms of Use Governing Terms for nvidia/omni embed nemotron 3b model: NVIDIA OneWay Noncommercial License. ADDITIONAL INFORMATION: Qwen RESEARCH LICENSE AGREEMENT This project will download and install additional third party open source s…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy