WebOrganizer/Corpus 200B [Paper] [Website] [GitHub] This dataset is a pre processed version of the 1b 1x CommonCrawl pool from DataComps LM cleaned with (1) RefinedWeb filters and (2) BFF deduplication. We provide the resulting 200B token corpus annotated with two quality scores, WebOrganizer domains, and k means scores. Download the dataset by cloning the repository with Git LFS instead of HuggingFace's load dataset() . The dataset has the following folder structure: We also include statistics about the presence and co occurence of domains in the domain statistics/ folder, computed with the domain statistics.py script. Citation If you make use of this pre processed corpus in your work, please cite:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy