wikipedia paragraphs wikipedia paragraphs is a dataset generated from Wikipedia, designed for natural language processing (NLP) research. Each entry contains cleaned paragraph text and Wikilink information extracted from a Wikipedia page, along with useful metadata such as categories, templates, and the associated Wikidata QID. Dataset structure Configurations The dataset is organized into multiple configurations, such as enwiki 20260301 v1.2.0 . Each configuration corresponds to a specific Wikipedia dump from which the data was sourced. The version number denotes the version of the scripts used to process the data. Data instances Each instance in the dataset represents a single Wikipedia page, which can be an article , category , or template . The following example from the enwiki 20260301 v1.2.0 configuration shows the data generated from the English Wikipedia article on Éclair. Data Fields id ( string ): A unique ID for the instance, composed of lang , page id and revision id lang ( string ): The Wikipedia edition code of the page. page id ( int64 ): The page ID of the page. revision id ( int64 ): The revision ID of the page. page type ( string ): The page type, one of: "article…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy