Public Dataset for AA Omniscience: Evaluating Cross Domain Knowledge Reliability in Large Language Models AA Omniscience Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies. Leaderboard and detailed results Paper Introduction We introduce AA Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly abstain when its knowledge is insufficient. AA Omniscience is characterized by its penalty for incorrect guesses, distinct from both accuracy (number of questions answered correctly) and hallucination rate (proportion of incorrect guesses when model does not know the answer), making it extremely relevant for users to choose a model for their next domain specific task. Dataset description The full dataset comprises 6,000 total questions, split across economically significant domains. Questions are created using a question generation agent, which derives questions from authoritative sources and filters them based on similarity, difficulty, and ambiguity. As a result, AA Omniscience can easily be scaled across more domains and progre…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy