Dataset Card for MegaWika Dataset Description Homepage: HuggingFace Repository: HuggingFace Paper: [Coming soon] Leaderboard: [Coming soon] Point of Contact: Samuel Barham Dataset Summary MegaWika is a multi and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non English language, an automated English translation is provided. Furthermore, nearly 130 million English question/answer pairs were extracted from the passages, and FrameNet events occurring in the passages are detected using the LOME FrameNet parser. Dataset Creation The pipeline through which MegaWika was created is complex, and is described in more detail in the paper (linked above), but the following diagram illustrates the basic approach. Supported Tasks and Leaderboards MegaWika is meant to support research across a variety of tasks, including report generation, summarization, information retrieval, question answering, etc. Languages MegaWika is divided by Wikipedia language. Ther…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy