From April 23 to April 27, 2026, Zhaozhi Wang attended the International Conference on Learning Representations (ICLR 2026) in Rio de Janeiro, Brazil. During the conference, he presented our paper, VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning.
In this work, we identify that multimodal large language models often struggle with visual-spatial reasoning because visual tokens are overshadowed by language tokens in attention. To address this issue, we propose VideoAnchor, a plug-and-play module that reinforces shared visual cues across frames without retraining, leading to more coherent visual grounding and stronger performance on spatial reasoning benchmarks.
Congratulations to Zhaozhi and the group on this exciting presentation at ICLR 2026!
Previous post
The skeptic’s guide to generative AI assisted coding
Next post
Tong Zhang and Zhaozhi Wang joined CICV 2026