Skip to main navigation Skip to search Skip to main content

DGAT: Dynamic Gaussian Attenuate Transformer for Remote Sensing Image Change Captioning

  • Xi'an Jiaotong University
  • State Grid Corporation of China
  • National Key Laboratory of Radar Detection and Sensing
  • First Institute of Photogrammetry and Remote Sensing

Research output: Contribution to journalArticlepeer-review

1 Scopus citations

Abstract

The remote sensing image change captioning (RSICC) technique is designed to enhance geospatial analysis by generating semantic descriptions of differences observed in bi-temporal remote sensing imagery (RSI). Although Transformer-based methods have achieved significant advancements in this field, their standard global attention mechanism allows pixels to pay equal attention to all spatial positions, which lacks explicit prior knowledge of spatial proximity correlation, making the model unable to strengthen local associations and weaken long-distance connections. Furthermore, most existing methods rely on global average pooling when reinforcing channelwise representations, which often dilutes the features of key changed regions by the largely unchanged background, resulting in the inability of feature compression to focus on the critical positions. To address the aforementioned challenges, this article proposes a dynamic Gaussian attenuate Transformer (DGAT), which innovatively introduces a dynamic Gaussian attenuation (DGA) mechanism to model an attenuation law between visual tokens from two perspectives based on Euclidean distance. At the spatial level, DGA constrains visual attention through distance-related Gaussian attenuation, allowing it to prioritize attention on adjacent regions, thereby enhancing the detection of changes in local continuity; at the channel level, DGA identifies the core area based on the visual attention kernel and uses the joint optimization of dynamic Gaussian weighted pooling (DGWP) and channel modulation to focus on features at key positions, effectively enhancing the expression of important channels. The experimental results on three benchmark RSICC datasets demonstrate that our proposed DGAT achieves significantly superior results, verifying the effectiveness of the DGA mechanism.

Original languageEnglish
Article number4424316
JournalIEEE Transactions on Geoscience and Remote Sensing
Volume63
DOIs
StatePublished - 2025

Keywords

  • Attenuate mechanism
  • change captioning
  • channel attention
  • remote sensing
  • Transformer

Fingerprint

Dive into the research topics of 'DGAT: Dynamic Gaussian Attenuate Transformer for Remote Sensing Image Change Captioning'. Together they form a unique fingerprint.

Cite this