MMStar (Are We on the Right Way for Evaluating Large Vision Language Models?) 🌐 Homepage 🤗 Dataset 🤗 Paper 📖 arXiv GitHub Dataset Details As shown in the figure below, existing benchmarks lack consideration of the vision dependency of evaluation samples and potential data leakage from LLMs' and LVLMs' training data. Therefore, we introduce MMStar: an elite vision indispensible multi modal benchmark, aiming to ensure each curated sample exhibits visual dependency , minimal data leakage , and requires advanced multi modal capabilities . 🎯 We have released a full set comprising 1500 offline evaluating samples. After applying the coarse filter process and manual review, we narrow down from a total of 22,401 samples to 11,607 candidate samples and finally select 1,500 high quality samples to construct our MMStar benchmark. In MMStar, we display 6 core capabilities in the inner ring, with 18 detailed axes presented in the outer ring. The middle ring showcases the number of samples for each detailed dimension. Each core capability contains a meticulously balanced 250 samples . We further ensure a relatively even distribution across the 18 detailed axes. 🏆 Mini Leaderboard We show a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy