Skip to main navigation Skip to search Skip to main content

RAMP: Iterative Refinement and Adaptive Multi-granularity Perception for embodied dialog localization

  • Xi'an Jiaotong University
  • Amazon.com, Inc.
  • University of Illinois at Chicago

Research output: Contribution to journalArticlepeer-review

Abstract

Embodied dialogue localization aims to determine a target location on a given 2D map based on dialogues. This task is critical for various real-world applications, such as emergency search and rescue, where precise localization is essential. The model must focus on fine features of the 2D map, guided by multi-round natural language dialogues. While previous research has yielded satisfactory results in a rough range, practical applications demand more precise localization to facilitate navigation or object manipulation by embodied agents. To address this challenge, we introduce an Iterative Refinement and Adaptive Multi-granularity Perception network, namely RAMP, which aims to iteratively refine the target location while enhancing the interaction of features at different granularities. Experimental results on the WAY dataset show that our method outperforms the state-of-the-art methods in both single-shot (+29.34%) and multi-shot (+36.85%) settings. These results highlight the superior performance of RAMP and its significant advancement over existing models.

Original languageEnglish
Article number114282
JournalPattern Recognition
Volume180
DOIs
StatePublished - Dec 2026

Keywords

  • Embodied dialogue localization
  • Hybrid granularity awareness
  • Masked attention

Fingerprint

Dive into the research topics of 'RAMP: Iterative Refinement and Adaptive Multi-granularity Perception for embodied dialog localization'. Together they form a unique fingerprint.

Cite this