Dataset Card for "wiki40b" Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: https://research.google/pubs/pub49029/ Repository: More Information Needed Paper: More Information Needed Point of Contact: More Information Needed Size of downloaded dataset files: 0.00 MB Size of the generated dataset: 10.47 GB Total amount of disk used: 10.47 GB Dataset Summary Clean up text for 40+ Wikipedia languages editions of pages correspond to entities. The datasets have train/dev/test splits per language. The dataset is cleaned up by page filtering to remove disambiguation pages, redirect pages, deleted pages, and non entity pages. Each example contains the wikidata id of the entity, and the full Wikipedia article after page processing that removes non content sections and…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy