Accurate and robust spatial understanding is crucial for the control and safety of embodied AI and robotic systems. Yet current evaluation practices remain fragmented, where different tasks adopt distinct benchmarks and datasets. This impedes a comprehensive evaluation of spatial understanding models from a holistic scene-understanding perspective and limits fine-grained analysis of their performance under challenging out-of-distribution conditions. In this work we present Unreal3DSpace, a unified and controllable framework for comprehensive evaluation of spatial understanding. Our Unreal3DSpace features controllable and procedural generation of diverse and high-fidelity 3D scenes and the availability of comprehensive 2D and 3D ground-truth annotations. This enables the evaluation of a wide range of spatial understanding models within a single, consistent framework and allows systematic studies of model robustness under varying degrees of scene- and object-level distribution shifts. Moreover, the comprehensive 3D ground truths also support analysis of the inference trajectories of spatial reasoning models, allowing us to pinpoint the underlying factors that lead to failure cases. Experimental results on our Unreal3DSpace highlight key limitations of current spatial understanding models, particularly in handling distant objects, partial occlusion, and questions with multiple seemingly plausible answers. We further introduce an agentic model that integrates multiple spatial understanding modules to infer spatial relationships, which pinpoints the primary bottlenecks contributing to errors in final predictions. Overall, evaluation results on our Unreal3DSpace provide valuable insights into the weaknesses of existing spatial understanding models and inform directions for future research in this area. Our code and data is available from our project page.
IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2026-06-03
2026-06-26