Ovis2 4B It is recommended to use the latest version: Ovis2.5. Introduction GitHub Paper We are pleased to announce the release of Ovis2 , our latest advancement in multi modal large language models (MLLMs). Ovis2 inherits the innovative architectural design of the Ovis series, aimed at structurally aligning visual and textual embeddings. As the successor to Ovis1.6, Ovis2 incorporates significant improvements in both dataset curation and training methodologies. Key Features : Small Model Performance : Optimized training strategies enable small scale models to achieve higher capability density, demonstrating cross tier leading advantages. Enhanced Reasoning Capabilities : Significantly strengthens Chain of Thought (CoT) reasoning abilities through the combination of instruction tuning and preference learning. Video and Multi Image Processing : Video and multi image data are incorporated into training to enhance the ability to handle complex visual information across frames and images. Multilingual Support and OCR : Enhances multilingual OCR beyond English and Chinese and improves structured data extraction from complex visual elements like tables and charts. Model Zoo Ovis MLLMs Vi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy