CC OCR V2: Benchmarking Large Multimodal Models for Literacy in Real world Document Processing Dataset Summary CC OCR V2 is a comprehensive and challenging OCR benchmark tailored to real world document processing. It focuses on practical enterprise document processing tasks and incorporates hard and corner cases that are critical yet underrepresented in prior benchmarks. The dataset comprises 7,093 high difficulty samples covering 5 major OCR centric tracks: Text Recognition, Document Parsing, Document Grounding, Key Information Extraction, and Document Question Answering. Dataset Structure The dataset is structured hierarchically by task and sub task . Below is the statistical breakdown of the dataset: Task Sub task Samples : : : Extraction business transactions 340 public services 369 regulated records 300 Grounding object grounding 734 text grounding 734 Parsing complex table parsing 300 formula parsing 100 general documents parsing 300 info board parsing 26 molecular parsing 100 QA blueprint qa 100 dashboards fact qa 400 dashboards numeric qa 500 financial documents qa 1000 Recognition multi lingual recognition 640 natural scene recognition 1150 Total 7093 Data Instances Each s…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy