Dataset Card for Wikimedia Structured Wikipedia Dataset Description Homepage: https://enterprise.wikimedia.com/ Point of Contact: Stephanie Delbecque Total size: 44.42 GiB Quick Links Wikimedia Enterprise Structured Contents Documentation Data Dictionary Wikimedia Attribution Framework Meta Wiki Discussion Dataset Summary Pre parsed English and French Wikipedia articles, extracted using the Wikimedia Enterprise Snapshot API. This dataset contains all articles of the English and French language editions of Wikipedia, pre parsed and output as structured data with a consistent schema. The dataset is provided in Parquet format, optimized for high performance analytical queries and efficient storage. This version uses a unified, pinned schema across all files, making it compatible with DuckDB, pandas, Polars, and Apache Spark out of the box. New in this dataset: Parsed references and citations, connecting Wikipedia's knowledge with its sources of truth. Parsed tables, one of the most information heavy sections of Wikipedia pages. Credibility signals, for example referenceneed and referencerisk , signaling where information may not be sufficiently backed by sources. Improvements to list…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy