PP DocLayout plus L Introduction A higher precision layout area localization model trained on a self built dataset containing Chinese and English papers, PPT, multi layout magazines, contracts, books, exams, ancient books and research reports using RT DETR L. The layout detection model includes 20 common categories: document title, paragraph title, text, page number, abstract, table, references, footnotes, header, footer, algorithm, formula, formula number, image, table, seal, figure table title, chart, and sidebar text and lists of references. The key metrics are as follow: Model mAP(0.5) (%) PP DocLayout plus L 83.2 Note : the evaluation set of the above precision indicators is the self built version sub area detection data set, including Chinese and English papers, magazines, newspapers, research reports PPT、 1000 document type pictures such as test papers and textbooks. Quick Start Installation 1. PaddlePaddle Please refer to the following commands to install PaddlePaddle using pip: For details about PaddlePaddle installation, please refer to the PaddlePaddle official website. 2. PaddleOCR Install the latest version of the PaddleOCR inference package from PyPI: Model Usage You…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy