Dataset Card for "pg19" Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: https://github.com/deepmind/pg19 Repository: More Information Needed Paper: Compressive Transformers for Long Range Sequence Modelling Point of Contact: More Information Needed Size of downloaded dataset files: 11.74 GB Size of the generated dataset: 11.51 GB Total amount of disk used: 23.25 GB Dataset Summary This repository contains the PG 19 language modeling benchmark. It includes a set of books extracted from the Project Gutenberg books library, that were published before 1919. It also contains metadata of book titles and publication dates. PG 19 is over double the size of the Billion Word benchmark and contains documents that are 20X longer, on average, than the WikiText long range…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy