ETRI-Knowledge Sharing Plaform

KOREAN
논문 검색
Type SCI
Year ~ Keyword

Detail

Conference Paper Can VLMs Handle Multi-hop Compositional Spatial Reasoning?
Cited - time in scopus Share share facebook twitter linkedin kakaostory
Authors
Youngwan Lee, Soojin Jang, Yoorhim Cho, Seunghwan Lee, Yong-Ju Lee, Sung Ju Hwang
Issue Date
2026-06
Citation
Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) 2026, pp.1-5
Publisher
Computer Vision Foundation
Language
English
Type
Conference Paper
Abstract
Spatial reasoning is a critical capability for Vi- sion–Language Models (VLMs), particularly when deployed as Vision–Language–Action (VLA) agents in real-world environments. However, existing benchmarks predominantly focus on simple, single-hop spatial ques- tions, falling short of capturing the multi-hop reasoning and precise visual grounding required in practical scenar- ios. To address this gap, we introduce MultihopSpatial, a benchmark designed for multi-hop compositional spatial reasoning with 1–3 hop questions across ego- and exo- centric perspectives. Through extensive evaluation of 30 state-of-the-art VLMs, we demonstrate that compositional spatial reasoning remains a significant challenge for current VLMs.
KSP Keywords
Extensive evaluation, Language Models, Multi-Hop, Real-world, Spatial reasoning, single-hop, state-of-The-Art