MMLU Multi Prompt Evaluation Data (correctness scores) Overview This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi prompt evaluation of LLMs." arXiv preprint arXiv:2405.17202 (2024). Dataset Details The MMLU benchmark comprises 57 diverse subjects and approximately 14,000 examples. It is a multiple choice question answering benchmark that tests the performance of LLMs across a wide range of topics. The data includes evaluation for 15 different SOTA LLMs and 100 different prompt templates. In this dataset, each row represents a different prompt template while each column represents each MMLU example. If you are interested in the full data, including used prompts and examples text, please see it here. The data from a specific subject can be downloaded using If you want to download the full data you can loop over all subjects Citing @artic…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy