TY - GEN
T1 - RSG-VLN
T2 - 28th International Conference on Intelligent Transportation Systems, ITSC 2025
AU - Chen, Liming
AU - Guo, Zhongyu
AU - Wang, Ningnan
AU - Ji, Haoxuan
AU - Chen, Weihuang
AU - Wang, Yitian
AU - Sun, Hongbin
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Vision-Language Navigation (VLN) empowers autonomous agents to navigate unseen environments through natural language instructions and visual observations. Despite advancements driven by foundational models, critical challenges persist, including inefficient trajectories due to static semantic maps, error propagation in multi-step instruction decomposition, and computational bottlenecks in real-world deployment. This paper presents Relevance Semantic Map-Guided VLN (RSG-VLN), a new framework that addresses these limitations by dynamically aligning semantic scene understanding with task objectives. Specially, RSG-VLN introduces relevance semantic maps, leveraging Pointwise Mutual Information to quantify contextual associations between objects and task goals, enabling agents to prioritize actions with high semantic-task relevance. Additionally, large language models is employed to decompose long-horizon instructions into temporally sequenced subtasks, each mapped to localized scene targets. A hybrid navigation strategy integrates global path planning-connecting subtask waypoints-with local obstacle avoidance using classical exploration methods, ensuring robustness in complex layouts. Extensive experiments on Habitat-Matterport 3D and Matterport 3D datasets demonstrate that RSG-VLN achieves state-of-the-art performance in various VLN tasks. These advancements underscore the potential of RSG-VLN as a scalable solution for real-world applications that require precise, context-aware navigation.
AB - Vision-Language Navigation (VLN) empowers autonomous agents to navigate unseen environments through natural language instructions and visual observations. Despite advancements driven by foundational models, critical challenges persist, including inefficient trajectories due to static semantic maps, error propagation in multi-step instruction decomposition, and computational bottlenecks in real-world deployment. This paper presents Relevance Semantic Map-Guided VLN (RSG-VLN), a new framework that addresses these limitations by dynamically aligning semantic scene understanding with task objectives. Specially, RSG-VLN introduces relevance semantic maps, leveraging Pointwise Mutual Information to quantify contextual associations between objects and task goals, enabling agents to prioritize actions with high semantic-task relevance. Additionally, large language models is employed to decompose long-horizon instructions into temporally sequenced subtasks, each mapped to localized scene targets. A hybrid navigation strategy integrates global path planning-connecting subtask waypoints-with local obstacle avoidance using classical exploration methods, ensuring robustness in complex layouts. Extensive experiments on Habitat-Matterport 3D and Matterport 3D datasets demonstrate that RSG-VLN achieves state-of-the-art performance in various VLN tasks. These advancements underscore the potential of RSG-VLN as a scalable solution for real-world applications that require precise, context-aware navigation.
KW - large language model
KW - semantic map
KW - visual language navigation
UR - https://www.scopus.com/pages/publications/105036966677
U2 - 10.1109/ITSC60802.2025.11423232
DO - 10.1109/ITSC60802.2025.11423232
M3 - 会议稿件
AN - SCOPUS:105036966677
T3 - IEEE Conference on Intelligent Transportation Systems, Proceedings, ITSC
SP - 4356
EP - 4361
BT - IEEE Intelligent Transportation Systems Conference, ITSC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 18 November 2025 through 21 November 2025
ER -