ARTE Group

2026

Mingrui Wu presented OpenBench at CVPR 2026

conference CVPR MLLM spatial reasoning OpenBench

From June 3 to June 7, 2026, Mingrui Wu attended the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026) at the Colorado Convention Center in Denver, Colorado. During the conference, he presented our paper, From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs (arXiv).

In this work, we introduce OpenBench, a benchmark for evaluating the spatial reasoning ability of multimodal large language models in realistic open-world scenes. Built from pedestrian-perspective multimodal data with metric 3D cues, OpenBench covers question types ranging from qualitative relations to quantitative and kinematic reasoning. Our results show that performance gains reported in indoor settings do not transfer directly to open-world scenarios, highlighting an important gap in grounded spatial intelligence.

Congratulations to Mingrui and the team on sharing this work with the CVPR 2026 community!


Yanhao Wu introduces AlignDrive for coordinated autonomous-driving planning

autonomous driving end-to-end planning arXiv AlignDrive

Congratulations to the Team on Achieving State-of-the-Art in Autonomous Driving! We are thrilled to announce that our recent joint work with Horizon Robotics, led by Yanhao Wu, has achieved Top-1 performance on both the Bench2drive and the newly released Fail2Drive official leaderboard. Notably, our approach demonstrated exceptional resilience, proving to have the best generalization ability among competing methods.

Discover how we did it in our paper: AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving (arXiv). For a deeper dive into the methodology and to view our qualitative results, please visit our official project homepage.

AlignDrive takes aim at a subtle but important planning mismatch in autonomous driving: where to go and how fast to move should be decided together. The method conditions longitudinal prediction on the planned drive path, so speed reasoning becomes tied to the vehicle’s lateral choices and surrounding-agent interactions. It also uses planning-focused augmentation to expose the model to rare safety-critical cases, improving robustness on challenging closed-loop benchmarks.


Tong Zhang and Zhaozhi Wang joined CICV 2026

conference CICV autonomous driving world models

Prof. Tong Zhang, together with team member Zhaozhi Wang, was invited to participate in CICV 2026, the 13th Congress of Intelligent and Connected Vehicles Technology, held in Shanghai from May 21 to May 22, 2026.

During the A2 workshop, World Model: A New Paradigm for Autonomous Driving Research from High-Quality Data Generation to Reinforcement Learning Training, Zhaozhi delivered a talk on behalf of Prof. Zhang’s team. The presentation, titled Exploring Safety for Autonomous Driving: From Scene Generation to Planning, discussed how scenario generation and planning methods can support safer autonomous-driving systems. More information about the conference is available on the CICV official website.

Congratulations to Prof. Zhang, our team member Zhaozhi, and the team on sharing our autonomous-driving research with the CICV community.


Zhaozhi Wang attended ICLR 2026 in Rio de Janeiro

conference ICLR visual-spatial reasoning VideoAnchor

From April 23 to April 27, 2026, Zhaozhi Wang attended the International Conference on Learning Representations (ICLR 2026) in Rio de Janeiro, Brazil. During the conference, he presented our paper, VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning.

In this work, we identify that multimodal large language models often struggle with visual-spatial reasoning because visual tokens are overshadowed by language tokens in attention. To address this issue, we propose VideoAnchor, a plug-and-play module that reinforces shared visual cues across frames without retraining, leading to more coherent visual grounding and stronger performance on spatial reasoning benchmarks.

Congratulations to Zhaozhi and the group on this exciting presentation at ICLR 2026!