SkyPile 150B Dataset Summary SkyPile 150B is a comprehensive, large scale Chinese dataset specifically designed for the pre training of large language models. It is derived from a broad array of publicly accessible Chinese Internet web pages. Rigorous filtering, extensive deduplication, and thorough sensitive data filtering have been employed to ensure its quality. Furthermore, we have utilized advanced tools such as fastText and BERT to filter out low quality data. The publicly accessible portion of the SkyPile 150B dataset encompasses approximately 233 million unique web pages, each containing an average of over 1,000 Chinese characters. In total, the dataset includes approximately 150 billion tokens and 620 gigabytes of plain text data. Language The SkyPile 150B dataset is exclusively composed of Chinese data. Data Field Explanation text: the processed and cleaned text extracted from each page. Dataset Safety We utilized more than 200w rules and the BERT base model to determine the sensitive data present in the dataset, and subsequently removed any harmful entries we detect. Sensitive Information and Bias Despite our best efforts, SkyPile 150B, given its construction from public…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy