✨ Introduction MinerU Popo is a lightweight and universal framework for POst Processing OCR outputs, bridging the gap between page level OCR parsing and document level semantic structure. It construct document tree structure based on with a 4B post processing model performing four subtasks: table truncation analysis, text truncation analysis, title hierarchy analysis, and image text association analysis. We handle the challenges of cross page geometric discontinuity, redundant document parsing and scalability to long documents via: Task Oriented Data Engine : Generate representative training data and simplify the task specific input. Dynamic Chunking and Synchronization : Process long document by dynamic chunks and reduce deviations across chunks to preserve global consistency. Document Enrichment : Structurally construct a tree, semantically generate summaries and split long section nodes. 📊 Performance Better Hierarchy (TEDS) after Post Processing Basic OCR Before After : : : : : : MinerU 53.7 90.6 MonkeyOCR 48.9 87.4 Dolphin 60.4 83.5 PaddleOCR 59.3 82.6 GLM OCR 53.5 81.8 Advantages Compared to Directly Using Pre trained Model Model TEDS Doc/s : : : : : : MinerU Popo 90.6 0.37…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy