ColNomic Embed Multimodal 3B: State of the Art Visual Document Retrieval colnomic embed multimodal 3b is a multi vector state of the art multimodal embedding model that excels at visual document retrieval tasks: High Performance : Achieves 61.2 NDCG@5 on Vidore v2, outperforming all other models except ColNomic Embed Multimodal 7B Unified Text Image Encoding : Directly encodes interleaved text and images without complex preprocessing Advanced Architecture : 3B parameter multimodal embedding model Open Weights : Model weights available for research use Performance Model Avg. ESG Restaurant Human Econ Macro Multi. AXA Multi. MIT Bio ESG Restaurant Synth. ESG Restaurant Synth. Multi. MIT Bio Multi. AXA Econ. Macro ColNomic Embed Multimodal 7B 62.7 73.9 54.7 61.3 66.1 57.3 56.7 64.2 68.3 61.6 ColNomic Embed Multimodal 3B 61.2 65.8 55.4 61.0 63.5 56.6 57.2 62.5 68.8 60.2 T Systems ColQwen2.5 3B 59.9 72.1 51.2 60.0 65.3 51.7 53.3 61.7 69.3 54.8 Nomic Embed Multimodal 7B 59.7 65.7 57.7 59.3 64.0 49.2 51.9 61.2 66.3 63.1 GME Qwen2 7B 59.0 65.8 56.2 55.4 64.0 54.3 56.7 55.1 60.7 62.9 Nomic Embed Multimodal 3B 58.8 59.8 57.5 58.8 62.5 49.4 49.4 58.6 69.6 63.5 Llama Index vdr 2b multi v1 58.4…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy