Dataset Description Repository: https://github.com/cisnlp/GlotCC Paper: https://arxiv.org/abs/2410.23825 Point of Contact: amir@cis.lmu.de Dataset Summary GlotCC V1.0 is a document level, general domain dataset derived from CommonCrawl, covering more than 1000 languages. It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl. We release our pipeline as open source at https://github.com/cisnlp/GlotCC. List of Languages: See https://datasets server.huggingface.co/splits?dataset=cis lmu/GlotCC V1 to get the list of splits available. Usage (Huggingface Hub Recommended) Replace bal Arab with your specific language. For faster downloads, make sure to pip install huggingface hub[hf transfer] and set the environment variable HF HUB ENABLE HF TRANSFER =1. Then you can load it with any library that supports Parquet files, such as Pandas: Usage (Huggingface datasets) Usage (Huggingface datasets streaming=True) Usage (direct download) If you prefer not to use the Hugging Face datasets or hub you can download it directly. For example, to download the first file of bal Arab : Additional Information The dataset is currently heavily under audit and changes ac…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy