EN 中文 SenseNova SI: Scaling Spatial Intelligence with Multimodal Foundation Models Overview Despite remarkable progress, multimodal foundation models still exhibit surprising deficiencies in spatial intelligence. In this work, we explore scaling up multimodal foundation models to cultivate spatial intelligence within the SenseNova SI family , built upon established multimodal foundations including visual understanding models (i.e., Qwen3 VL and InternVL3) and unified understanding and generation models (i.e., Bagel). We take a principled approach to constructing high performing and robust spatial intelligence by systematically curating SenseNova SI 8M: eight million diverse data samples under a rigorous taxonomy of spatial capabilities. SenseNova SI demonstrates unprecedented performance across a broad range of spatial intelligence benchmarks, while maintaining strong general multimodal understanding. More importantly, we analyze the impact of data scaling, discuss early signs of emergent generalization capabilities enabled by diverse data training, analyze the risk of overfitting and language shortcuts, present a preliminary study on spatial chain of thought reasoning, and validat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy