Dataset Card for "wikitext" Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: https://blog.einstein.ai/the wikitext long term dependency language modeling dataset/ Repository: More Information Needed Paper: Pointer Sentinel Mixture Models Point of Contact: Stephen Merity Size of downloaded dataset files: 391.41 MB Size of the generated dataset: 1.12 GB Total amount of disk used: 1.52 GB Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText 2 is over 2 times larger and WikiText 103 is over 11…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy