Model Card [\[๐ H2OVL Mississippi Paper\]](https://arxiv.org/abs/2410.13611) [\[๐ค HF Demo\]](https://huggingface.co/spaces/h2oai/h2ovl mississippi) [\[๐ Quick Start\]]( quick start) The H2OVL Mississippi 800M is a compact yet powerful vision language model from H2O.ai, featuring 0.8 billion parameters. Despite its small size, it delivers state of the art performance in text recognition, excelling in the Text Recognition segment of OCRBench and outperforming much larger models in this domain. Built upon the robust architecture of our H2O Danube language models, the Mississippi 800M extends their capabilities by seamlessly integrating vision and language tasks. Key Features: 0.8 Billion Parameters: Balance between performance and efficiency, making it suitable for OCR and document processing. Trained on 19 million image text pairs, with a focus on OCR, document comprehension, and chart, figure, and table interpretation, the model is optimized for superior OCR performance. Benchmarks Performance Comparison of Similar Sized Models Across Multiple Benchmarks OpenVLM Leaderboard Models Params (B) Avg. Score MMBench MMStar MMMU VAL Math Vista Hallusion AI2D TEST OCRBench MMVet Qwen2 VLโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy