CulturaY: A Large Cleaned Multilingual Dataset of 75 Languages Dataset Summary From the team that brought you CulturaX, we present CulturaY, another substantial multilingual dataset of 15TB (uncompressed)/3TB (zstd compressed) that applies the same dataset cleaning methodology to the HPLT v1.1 dataset. Please note that HPLT v1.2 has also been released and is an alternative verison with different cleaning methodolgies. This data was used in part to train our SOTA Vietnamese model: Vistral 7B Chat. Our annotations and arrangements are licensed under CC BY 4.0, and we make the data available for fair use machine learning research. But we make no claims as to the underlying copyrights of the work. This data was copied from the HPLT project, which in turn used the data from Common Crawl and the Internet Archive. Acknowledgement We thank our collaborators at UONLP The Natural Language Processing Group at the University of Oregon, and the computing resources of the managers of the Karolina Supercomputers. We also thank our friends at TurkuNLP for their support. Data Breakdown: There are 75 langauges, with the following breakdown: Code Language Documents Documents (%) Size (GB) : : : : : :…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy