Measuring Multimodal Mathematical Reasoning with the MATH Vision Dataset [π» Github] [π Homepage] [π Main Leaderboard ] [π Open Source Leaderboard ] [πΏ Wild Leaderboard ] [π Visualization] [π Paper] πΏ NEW: MATH Vision Wild MATH Vision Wild is a photographic, real world variant of MATH Vision. The same testmini problems are physically captured on printed paper, iPads, laptops, and projectors under varying lighting and angles β the conditions VLMs actually face when a user holds up a phone to a math problem. π¦ Dataset : MathLLMs/MathVision Wild Β· π Leaderboard : mathllm.github.io/mathvision/ wildleaderboard Key finding β almost every model regresses in the wild: Model MATH Vision (testmini) MATH Vision Wild Ξ : : : o4 mini π₯ 55.9 57.2 +2.33% (only model to improve) Gemini 2.5 Pro Preview 05 06 (thinking) 63.8 49.0 β23.20% Gemini 2.5 Flash Preview 05 20 57.9 48.0 β17.10% Doubao 1.5 thinking vision pro 57.9 45.7 β21.07% Gemini 2.5 Pro Preview 05 06 61.8 42.8 β30.74% GPT 4.1 40.5 35.5 β12.35% Qwen2.5 VL 72B Instruct 36.2 24.0 β33.70% Gemini 2.0 Flash 48.0 23.0 β52.08% Gemini 1.5 Pro 38.8 18.4 β52.58% Only o4 mini improves when problems are photographed; long reasoning models dβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy