RepLiQA Repository of Likely Question Answer for benchmarking NeurIPS Datasets presentation Dataset Summary RepLiQA is an evaluation dataset that contains Context Question Answer triplets, where contexts are non factual but natural looking documents about made up entities such as people or places that do not exist in reality. RepLiQA is human created, and designed to test for the ability of Large Language Models (LLMs) to find and use contextual information in provided documents. Unlike existing Question Answering datasets, the non factuality of RepLiQA makes it so that the performance of models is not confounded by the ability of LLMs to memorize facts from their training data: one can test with more confidence the ability of a model to leverage the provided context. Documents in RepLiQA comprise 17 topics or document categories: Company Policies ; Cybersecurity News ; Local Technology and Innovation ; Local Environmental Issues ; Regional Folklore and Myths ; Local Politics and Governance ; News Stories ; Local Economy and Market ; Local Education Systems ; Local Arts and Culture ; Local News ; Small and Medium Enterprises ; Incident Report ; Regional Cuisine and Recipes ; Neighb…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy