Arabic Books Dataset Summary The arabic books dataset contains 8,500 rows of text , each representing the full text of a single Arabic book. These texts were extracted using the arabic large nougat model, showcasing the model’s capabilities in Arabic OCR and text extraction. The dataset spans a total of 1.1 billion tokens , calculated using the GPT 4 tokenizer. This dataset is a testimony to the quality of the Arabic Nougat models and their effectiveness in extracting structured text from complex Arabic documents. Research Context The dataset is part of the Arabic Nougat research project and supports the findings of the research paper: Arabic Nougat: Fine Tuning Vision Transformers for Arabic OCR and Markdown Extraction . Purpose This dataset is intended to: Demonstrate the performance of Arabic Nougat models. Serve as a resource for developing Arabic language models and advancing research in Arabic NLP . Licensing This dataset is released under the GPL 3.0 License , enabling its open source availability for further research and development. Citation If you use this dataset, please cite the corresponding research paper:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy