ARTE Group

CVPR 2026

From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs

OpenBench: A benchmark for open-world spatial reasoning in multimodal large language models

Mingrui Wu Zhaozhi Wang Fangjinhua Wang Jiaolong Yang Marc Pollefeys Tong Zhang

CVPR ยท 2026

From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs teaser image

Abstract

We present OpenBench to evaluate how well multimodal large language models perform spatial reasoning in realistic open-world scenes. The benchmark is built from pedestrian-perspective multimodal data with metric 3D cues, and includes question types ranging from qualitative relations to quantitative and kinematic reasoning. Results show that performance gains reported in indoor settings do not transfer well to open-world scenes, highlighting a major gap in grounded spatial intelligence.

Citation

@inproceedings{wu2026indooropenworldrevealing,
      title={From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs},
      author={Mingrui Wu and Zhaozhi Wang and Fangjinhua Wang and Jiaolong Yang and Marc Pollefeys and Tong Zhang},
      booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
      year={2026},
      eprint={2512.19683},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2512.19683},
}