This is our basic reasoning model before enhanced by Reinforcement Learning. Multimodal LLM x Reasoning Model π π§ π After more than six months since creating the 5CD AI/LLaVA CoT o1 Instruct datasetβone of Hugging Faceβs most liked datasets of 2024 πβwe have just completed the "base" version of the Vintern Reasoning Model! This model can perform long and complex reasoning based on images, breaking down each reasoning step into multiple sub steps while keeping hallucinations under control. Despite the difficulty of balancing multiple tasks alongside reasoning, Vintern 3B R beta has outperformed all previous versions across various benchmarks! When should you choose Vintern 1B v3 5 vs Vintern 3B R beta? π€ Vintern 1B v3 5 : Fast β‘ and good for Vietnamese OCR with simple text formatting. π Highly reliable. β Vintern 3B R beta : Better for complex questions and complex structured doc image. ππ OCR performance on blurred or unclear text may be slightly affected due to our training focus on reasoning. ππ€ π The next step? Training and enhancing its reasoning ability by Reinforcement Learning! Benchmarks π Example 1: Example 2: Quickstart Here provides a code snippet to show youβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy