General OCR Theory: Towards OCR 2.0 via a Unified End to end Model 🔋Online Demo 🌟GitHub 📜Paper Haoran Wei , Chenglong Liu , Jinyue Chen, Jia Wang, Lingyu Kong, Yanming Xu, Zheng Ge, Liang Zhao, Jianjian Sun, Yuang Peng, Chunrui Han, Xiangyu Zhang Usage Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.10: More details about 'ocr type', 'ocr box', 'ocr color', and 'render' can be found at our GitHub. Our training codes are available at our GitHub. More Multimodal Projects 👏 Welcome to explore more multimodal projects of our team: Vary Fox OneChart Citation If you find our work helpful, please consider citing our papers 📝 and liking this project ❤️!
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy