LayoutLMv3 Microsoft Document AI GitHub Model description LayoutLMv3 is a pre trained multimodal Transformer for Document AI with unified text and image masking. The simple unified architecture and training objectives make LayoutLMv3 a general purpose pre trained model. For example, LayoutLMv3 can be fine tuned for both text centric tasks, including form understanding, receipt understanding, and document visual question answering, and image centric tasks such as document image classification and document layout analysis. LayoutLMv3: Pre training for Document AI with Unified Text and Image Masking Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, Furu Wei, Preprint 2022. Citation If you find LayoutLM useful in your research, please cite the following paper: License The content of this project itself is licensed under the Attribution NonCommercial ShareAlike 4.0 International (CC BY NC SA 4.0). Portions of the source code are based on the transformers project. Microsoft Open Source Code of Conduct
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy