We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy
This dataset is the result of processing all WARC files in the CCNews Corpus, from the beginning (2016) to June of 2024. The data has been cleaned and deduplicated, and language of articles have been detected and added. The process is similar to what HuggingFace's DataTrove does. Overall, it contains about 600 million news articles in more than 100 languages from all around the globe. For license information, please refer to CommonCrawl's Terms of Use. Sample Python code to explore this… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/ccnews.
No dataset card provided yet.
Mirrored from an external registry.
Last synced 6/13/2026
Preview not yet available for this dataset