Abstract
Embodied dialogue localization aims to determine a target location on a given 2D map based on dialogues. This task is critical for various real-world applications, such as emergency search and rescue, where precise localization is essential. The model must focus on fine features of the 2D map, guided by multi-round natural language dialogues. While previous research has yielded satisfactory results in a rough range, practical applications demand more precise localization to facilitate navigation or object manipulation by embodied agents. To address this challenge, we introduce an Iterative Refinement and Adaptive Multi-granularity Perception network, namely RAMP, which aims to iteratively refine the target location while enhancing the interaction of features at different granularities. Experimental results on the WAY dataset show that our method outperforms the state-of-the-art methods in both single-shot (+29.34%) and multi-shot (+36.85%) settings. These results highlight the superior performance of RAMP and its significant advancement over existing models.
| Original language | English |
|---|---|
| Article number | 114282 |
| Journal | Pattern Recognition |
| Volume | 180 |
| DOIs | |
| State | Published - Dec 2026 |
Keywords
- Embodied dialogue localization
- Hybrid granularity awareness
- Masked attention
Fingerprint
Dive into the research topics of 'RAMP: Iterative Refinement and Adaptive Multi-granularity Perception for embodied dialog localization'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver