Dataset Card for NewsWire Dataset Description Homepage: Dell Research homepage Repository: Github repository Paper: arxiv submission Point of Contact: Melissa Dell Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is divided into a distinct file (eg. 1952 data clean.json) Data Instances An example from the NewsWire dataset looks like: Data Fields year : year of article publication. dates : list of dates on which this article was published, as strings in the form mmm DD YYYY. byline : article byline, if any. article : article text. newspaper metadata : list of newspapers that carried the article. Each newspaper is represented as a list of dictionaries, where lccn is the newspaper's Library of Congress identifier, newspaper title is the name of the newspaper, and newspaper city and newspaper state give the location of t…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy