MedThinkVQA MedThinkVQA is an expert annotated benchmark for multi image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image , integrate cross view information , and perform differential diagnosis reasoning . Links GitHub: https://github.com/benluwang/MedThinkVQA Leaderboard: https://benluwang.github.io/MedThinkVQA/ Submission Guide: https://benluwang.github.io/MedThinkVQA/submit.html Benchmark at a Glance MedThinkVQA is designed to be hard for current multimodal models. Benchmark results show three consistent patterns: performance improves when models can use more images, when they are allowed stronger reasoning effort, and when larger multimodal variants are used within the same family. Benchmark observations. Top: current open and closed models still have substantial headroom on the test benchmark. Bottom left: using more images materially improves diagnostic reasoning. Bottom middle: stronger reasoning effort also improves performance. Bottom right: larger multimodal model variants in the same family tend to perform better on MedThinkVQA. Think w…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy