PP OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion Scale VLMs on OCR Tasks 🔥 Official Website 📝 Technical Report PP OCRv6 Overview PP OCRv6 is a lightweight OCR system that combines architectural innovation with data centric optimization. It redesigns the backbone, detection neck, and recognition neck around a unified MetaFormer style building block with structural reparameterization. Three model tiers (medium, small, tiny) share the same block primitives, covering deployment scenarios from server to edge. Key Features 1. Unified and Scalable Model Family: A three tier OCR model family spanning 1.5M to 34.5M parameters. PP OCRv6 medium achieves 86.2% detection Hmean and 83.2% recognition accuracy, outperforming PP OCRv5 server by +4.6% and +5.1% respectively. 2. Lightweight Architectural Innovations: (i) LCNetV4, a MetaFormer style lightweight backbone with structural reparameterization; (ii) RepLKFPN, a detection neck with dilated reparameterizable depthwise convolutions; (iii) EncoderWithLightSVTR, a recognition neck with local global attention and additive skip connections. 3. Multi Language and Scenario Support: Supports 50 languages and diverse industrial scenes (di…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy