Haon Chen/e5 omni 7B e5 omni 7B is a high performance omni modal embedding model built on top of Qwen2.5 Omni 7B. It produces a single, unified embedding space for text, images, audio, and video—making cross modal retrieval accurate and easy to use across a wide range of applications. Paper. 📝 Text 🖼️ Image 🎧 Audio 🎥 Video Experimental Results Our model achieves strong performance on MMEB V2 and AudioCaps benchmarks. Usage Using Sentence Transformers Install Sentence Transformers with the multimodal extras (for image, audio, and video support): 🎬 Video Retrieval 🎵 Audio Retrieval 📈 Image Document Retrieval (Image, Chart, PDF) 🌍 Multilingual Text Retrieval 🧩 Multimodal Inputs To embed a document that combines multiple modalities, pass a dict with any combination of "text" , "image" , "audio" , and "video" keys instead of a single path or string: Using Transformers The examples below are adapted from Tevatron. 🎬 Video Retrieval 🎵 Audio Retrieval 📈 Image Document Retrieval (Image, Chart, PDF) 🌍 Multilingual Text Retrieval Citation If you use this model in your research, please cite the associated paper.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy