OvisOCR2 Technical Report Online Demo Introduction We are pleased to announce the release of OvisOCR2, a compact 0.8B end to end model for page level document parsing. Given a document page image, OvisOCR2 generates a Markdown representation in natural reading order, covering text, formulas, tables, and visual regions. OvisOCR2 is developed by post training Qwen3.5 0.8B using a carefully designed data engine that combines real world and synthetic data, together with a multi stage training recipe integrating SFT, RL, and OPD. The model delivers strong document parsing performance while maintaining a small deployment footprint. OvisOCR2 achieves an overall score of 96.58 on OmniDocBench v1.6, establishing a new state of the art and becoming the first end to end model to top this leaderboard previously dominated by pipeline methods . On PureDocBench, OvisOCR2 also achieves the highest Avg3 score of 75.06. Performance Inference By default, parse removes HTML image tags for visual regions. To render Markdown with visual regions, set filter imgtags=False and save the Markdown file together with the referenced image crops as follows: Citation If you find OvisOCR2 useful, ple…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy