Wikitext Document Level This is a modified version of https://huggingface.co/datasets/wikitext that returns Wiki pages instead of Wiki text line by line. The original readme is contained below. Dataset Card for "wikitext" Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: https://blog.einstein.ai/the wikitext long term dependency language modeling dataset/ Repository: More Information Needed Paper: Pointer Sentinel Mixture Models Point of Contact: Stephen Merity Size of downloaded dataset files: 373.28 MB Size of the generated dataset: 1072.25 MB Total amount of disk used: 1445.53 MB Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The da…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy