CVPR 2026
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
OpenBench: A benchmark for open-world spatial reasoning in multimodal large language models
CVPR ยท 2026
Abstract
We present OpenBench to evaluate how well multimodal large language models perform spatial reasoning in realistic open-world scenes. The benchmark is built from pedestrian-perspective multimodal data with metric 3D cues, and includes question types ranging from qualitative relations to quantitative and kinematic reasoning. Results show that performance gains reported in indoor settings do not transfer well to open-world scenes, highlighting a major gap in grounded spatial intelligence.
Citation
@inproceedings{wu2026indooropenworldrevealing,
title={From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs},
author={Mingrui Wu and Zhaozhi Wang and Fangjinhua Wang and Jiaolong Yang and Marc Pollefeys and Tong Zhang},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2026},
eprint={2512.19683},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.19683},
}