Dataset Card for Wikipedia Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: https://dumps.wikimedia.org Repository: More Information Needed Paper: More Information Needed Point of Contact: More Information Needed Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). The articles are parsed using the mwparserfromhell tool, which can be installed with: Then, you can load any subset of Wikipedia per language and per date this way: [!TIP] You can specify num proc= in load dataset to generate the d…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy