摘要
Embodied dialogue localization aims to determine a target location on a given 2D map based on dialogues. This task is critical for various real-world applications, such as emergency search and rescue, where precise localization is essential. The model must focus on fine features of the 2D map, guided by multi-round natural language dialogues. While previous research has yielded satisfactory results in a rough range, practical applications demand more precise localization to facilitate navigation or object manipulation by embodied agents. To address this challenge, we introduce an Iterative Refinement and Adaptive Multi-granularity Perception network, namely RAMP, which aims to iteratively refine the target location while enhancing the interaction of features at different granularities. Experimental results on the WAY dataset show that our method outperforms the state-of-the-art methods in both single-shot (+29.34%) and multi-shot (+36.85%) settings. These results highlight the superior performance of RAMP and its significant advancement over existing models.
| 源语言 | 英语 |
|---|---|
| 期刊论文编号 | 114282 |
| 期刊 | Pattern Recognition |
| 卷 | 180 |
| DOI | |
| 出版状态 | 已出版 - 12月 2026 |
学术指纹
探究 'RAMP: Iterative Refinement and Adaptive Multi-granularity Perception for embodied dialog localization' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver