MMMU Pro (A More Robust Multi discipline Multimodal Understanding Benchmark) π Homepage π Leaderboard π€ Dataset π€ Paper π arXiv GitHub πNews π οΈ[2026 05 30] Fixed the option augmentation issue in Vision and Standard (10 options) settings. (validation Diagnostics and Laboratory Medicine 17) π οΈ[2025 03 08] Fixed mismatch between inner image labels and shuffled options in Vision and Standard (10 options) settings. (test Chemistry 5,94,147,216,314,345,354,461,560,570; test Materials 450; test Pharmacy 198; validation Chemistry 12,26,29; validation Materials 10,28; validation Psychology 1) π οΈ[2024 11 10] Added options to the Vision subset. π οΈ[2024 10 20] Uploaded Standard (4 options) cases. π₯[2024 09 05] Introducing MMMU Pro, a robust version of MMMU benchmark for multimodal AI evaluation! π Introduction MMMU Pro is an enhanced multimodal benchmark designed to rigorously assess the true understanding capabilities of advanced AI models across multiple modalities. It builds upon the original MMMU benchmark by introducing several key improvements that make it more challenging and realistic, ensuring that models are evaluated on their genuine ability to integrate and comprehend bβ¦
Runs entirely in your browser via DuckDB-Wasm β this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy