ARTE Group

Zhaozhi Wang attended ICLR 2026 in Rio de Janeiro

conference ICLR visual-spatial reasoning VideoAnchor

From April 23 to April 27, 2026, Zhaozhi Wang attended the International Conference on Learning Representations (ICLR 2026) in Rio de Janeiro, Brazil. During the conference, he presented our paper, VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning.

In this work, we identify that multimodal large language models often struggle with visual-spatial reasoning because visual tokens are overshadowed by language tokens in attention. To address this issue, we propose VideoAnchor, a plug-and-play module that reinforces shared visual cues across frames without retraining, leading to more coherent visual grounding and stronger performance on spatial reasoning benchmarks.

Congratulations to Zhaozhi and the group on this exciting presentation at ICLR 2026!