Abstract
The remote sensing image change captioning (RSICC) technique is designed to enhance geospatial analysis by generating semantic descriptions of differences observed in bi-temporal remote sensing imagery (RSI). Although Transformer-based methods have achieved significant advancements in this field, their standard global attention mechanism allows pixels to pay equal attention to all spatial positions, which lacks explicit prior knowledge of spatial proximity correlation, making the model unable to strengthen local associations and weaken long-distance connections. Furthermore, most existing methods rely on global average pooling when reinforcing channelwise representations, which often dilutes the features of key changed regions by the largely unchanged background, resulting in the inability of feature compression to focus on the critical positions. To address the aforementioned challenges, this article proposes a dynamic Gaussian attenuate Transformer (DGAT), which innovatively introduces a dynamic Gaussian attenuation (DGA) mechanism to model an attenuation law between visual tokens from two perspectives based on Euclidean distance. At the spatial level, DGA constrains visual attention through distance-related Gaussian attenuation, allowing it to prioritize attention on adjacent regions, thereby enhancing the detection of changes in local continuity; at the channel level, DGA identifies the core area based on the visual attention kernel and uses the joint optimization of dynamic Gaussian weighted pooling (DGWP) and channel modulation to focus on features at key positions, effectively enhancing the expression of important channels. The experimental results on three benchmark RSICC datasets demonstrate that our proposed DGAT achieves significantly superior results, verifying the effectiveness of the DGA mechanism.
| Original language | English |
|---|---|
| Article number | 4424316 |
| Journal | IEEE Transactions on Geoscience and Remote Sensing |
| Volume | 63 |
| DOIs | |
| State | Published - 2025 |
Keywords
- Attenuate mechanism
- change captioning
- channel attention
- remote sensing
- Transformer
Fingerprint
Dive into the research topics of 'DGAT: Dynamic Gaussian Attenuate Transformer for Remote Sensing Image Change Captioning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver