MulSeT: A Benchmark for Multi view Spatial Understanding Tasks Paper: Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Code: https://github.com/WanyueZhang ai/spatial understanding A high level overview of the MulSeT benchmark. The dataset challenges models to integrate information from two distinct viewpoints of a 3D scene to answer spatial reasoning questions. 📝 Dataset Summary MulSeT is a comprehensive benchmark designed to evaluate the multi view spatial understanding capabilities of Multimodal Large Language Models (MLLMs). The core challenge lies in integrating visual information from two different viewpoints of a 3D scene to answer complex spatial questions. All tasks are formulated as four option multiple choice questions, requiring models to perform sophisticated reasoning beyond simple object recognition. The dataset is synthetically generated using the AI2 THOR and replica cad , allowing for precise control over scene composition, object placement, and viewpoint selection. 🏆 News & Milestones 🎉 [2025 10 16] We've hit 100k+ downloads! A huge thank you to the entire community for your incredible interest and support. We'r…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy