Skip to main navigation Skip to search Skip to main content

AFSIFormer: Adaptive Frequency-Spatial Interaction Attention Mechanism for Aerial Image Semantic Segmentation

  • Jie Hui
  • , Wenyu Mi
  • , Jianji Wang
  • , Yuanyang Cao
  • , Ziyi Zhou
  • , Nanning Zheng
  • Xi'an Jiaotong University
  • Xi'an Jiaotong-Liverpool University

Research output: Contribution to journalArticlepeer-review

8 Scopus citations

Abstract

Aerial image semantic segmentation continues to face significant challenges in accurately capturing boundary textures. While convolutional neural networks (CNNs) and transformers are effective at modeling local features and long-range contextual dependencies, they often struggle with fine-grained boundary representation. In contrast, frequency-domain information offers complementary advantages by effectively representing periodic textures and structural edges. In this article, we propose a novel adaptive frequency-spatial interaction transformer (AFSIFormer) that follows a progressive learning strategy. First, a boundary-aware directional attention mechanism (BADAM) captures long-range dependencies across windows. Then, a local window attention mechanism (LWAM) refines contextual information within each window, enabling fine-grained local modeling under global guidance. Both BADAM and LWAM are built upon our designed adaptive frequency-spatial interaction attention (AFSIAttention) to capture contextual information across different windows. Unlike the existing frequency- and spatial-domain external coarse integration strategies, this mechanism utilizes a head-specific lightweight frequency projection network (HS-LFPN) to dynamically generate frequency-domain weights for each attention head. These frequency weights interact adaptively with spatial attention (SpatAttn) weights, facilitating frequency-guided spatial feature learning and internal integration of frequency and spatial information. Furthermore, we design a block-level residual coupling architecture that embeds AFSIFormer as a residual module within each convolutional stage, allowing continuous infusion of global and frequency-domain cues throughout the network. These collectively constitute the synergistic frequency-spatial network (SynFSNet), which achieves the state-of-the-art (SOTA) performance on three benchmark aerial image segmentation datasets. The code is available at https://github.com/Xinmu-Tantai/SynFSNet

Original languageEnglish
Article number5638419
JournalIEEE Transactions on Geoscience and Remote Sensing
Volume63
DOIs
StatePublished - 2025

Keywords

  • Aerial image segmentation
  • attention mechanism
  • frequency-spatial interaction
  • remote sensing
  • residual coupling architecture

Fingerprint

Dive into the research topics of 'AFSIFormer: Adaptive Frequency-Spatial Interaction Attention Mechanism for Aerial Image Semantic Segmentation'. Together they form a unique fingerprint.

Cite this