Dataset Card for PopQA Dataset Summary PopQA is a large scale open domain question answering (QA) dataset, consisting of 14k entity centric QA pairs. Each question is created by converting a knowledge tuple retrieved from Wikidata using a template. Each question come with the original subject entitiey , object entity and relationship type annotation, as well as Wikipedia monthly page views. Languages The dataset contains samples in English only. Dataset Structure Data Instances Size of downloaded dataset file: 5.2 MB Data Fields id : question id subj : subject entity name prop : relationship type obj : object entity name subj id : Wikidata ID of the subject entity prop id : Wikidata relationship type ID obj id : Wikidata ID of the object entity s aliases : aliases of the subject entity o aliases : aliases of the object entity s uri : Wikidata URI of the subject entity o uri : Wikidata URI of the object entity s wiki title : Wikipedia page title of the subject entity o wiki title : Wikipedia page title of the object entity s pop : Wikipedia monthly pageview of the subject entity o pop : Wikipedia monthly pageview of the object entity question : PopQA question possible answers : a li…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy