MMLU Pro Dataset MMLU Pro dataset is a more robust and challenging massive multi task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines. Github 🏆Leaderboard 📖Paper 🚀 What's New \[2026.03.11\] Added more cutting edge frontier models to the leaderboard, including the Claude 4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini 3.1 Pro, among others. Stay tuned for updated rankings and analysis. \[2026.01.18\] Fixed leading space issue in answer options (affected chemistry, physics, and other STEM subsets). This formatting inconsistency could have been exploited as a shortcut. Thanks to @giffmana and @fujikanaeda for identifying this. \[2025.10.25\] Posted a consolidated note on Health category issues and minor category updates (does not change overall micro averaged scores; may slightly affect per category metrics, mainly Health/Psychology). See details: https://huggingface.co/datasets/TIGER Lab/MMLU Pro/discussions/36. Special thanks to @mkieffer for the professional and meticulous review. \[2025.04.06\] We corrected 15 answers in medical domain based on the reco…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy