TomoroAI/tomoro colqwen3 embed 4b ⚡ Executive Summary TomoroAI/tomoro colqwen3 embed 4b is a state of the art ColPali style multimodal embedding model. It maps text queries, visual documents (images, PDFs) or short videos into aligned multi vector embeddings. Built by merging Qwen/Qwen3 VL 4B Instruct with Qwen/Qwen3 Embedding 4B , this model inherits robust text retrieval capabilities while preserving a full vision stack. It has been fine tuned on a curated mixture of VDR, ViDoRe ColPali Training, VisRAG Ret Train Synthetic data, and VisRAG Ret Train In domain data. It achieves SOTA or competitive performance across ViDoRe V1 V3 (English and Multilingual) while offering a significantly reduced embedding footprint compared to other full dim Colpali model alternatives. 🛠️ Model Specifications Feature Detail : : Architecture Qwen3 VL 4B (Encoder only variant) + 320 dim Projection Head Methodology ColPali style Late Interaction (MaxSim scoring) Token Budget Up to 1,280 visual tokens per page or 5120 visual tokens per video (text prompts constrained only by the base context window) Context Window 32k (inherited from base), typical usage < 2k tokens Output Multi vector (Seq Len × 320),…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy