Skip to main navigation Skip to search Skip to main content

引入全局感知与细节增强的非对称遥感建筑物分割网络

Translated title of the contribution: Global perception and detail enhancement network for building segmentation in remote sensing images
  • Shengjun Xu
  • , Yurui Liu
  • , Erhu Liu
  • , Jun Liu
  • , Ya Shi
  • , Xiaohan Li
  • Xi'an University of Architecture and Technology
  • Xi'an Key Laboratory of Intelligent Technology for Building Manufacturing

Research output: Contribution to journalArticlepeer-review

Abstract

Objective Remote sensing images are a type of earth observation data with wide coverage, rich spectral information, and variable target structures. The advancement of computer technology has steadily increased the demand for accurate and efficient extraction of buildings across diverse domains and industries. Meanwhile, the application prospects of semantic segmentation techniques in remote sensing image have progressively demonstrated substantial practical significance. By utilizing the semantic segmentation technology for remote sensing images, detailed information such as the spatial distribution and density of buildings and other infrastructures can be efficiently extracted. This information will play a crucial role in land surveying, urban planning, and post-disaster assessments. However, this advancement has simultaneously increased the complexity of semantic segmentation for buildings in remote sensing images. Consequently, the challenge of efficiently and accurately extracting building information from high-resolution imagery has emerged as a pivotal concern in the field of semantic segmentation of remote sensing images, which demands urgent attention and resolution. In recent years, deep learning has notably advanced in the field of semantic segmentation of remote sensing images. These advancements are due to its ability to learn any data distribution without requiring prior statistical knowledge of the input data, its capacity for self-learning target features, and its strong generalization capabilities. However, the process of semantic segmentation for remote sensing images of buildings faces substantial obstacles, which are primarily due to robust interferences such as varying lighting conditions, seasonal changes, and complex background information, as well as the intricate architectural structures and edge details of the buildings themselves. To address these challenges, this study proposes a global perception and detail enhancement asymmetric-UNet(GPDEA-UNet)network for building semantic segmentation in remote sensing images. Method First, the proposed network using UNet architecture constructs a feature encoder module based on the selective state space module. This module is specifically designed to meticulously extract the texture, boundary, and deep semantic features of buildings in remote sensing images. It leverages the visual state space as its fundamental building block and incorporates dynamic convolution decomposition(DCD)to significantly enhance the extraction of intricate features and context information in the remote sensing images while effectively reducing computational overhead. Second, a multi-scale dual cross-attention(MDCA)module is introduced to further broaden the global receptive field of the network and tackle the semantic discrepancy challenges posed by the codec during skip connections. MDCA represents an advanced attention-weighting mechanism that harmoniously integrates cross-channel attention and cross-spatial attention. This module substantially enhances the capability of the network to extract and fuse feature information pertinent to the region and boundary of the segmented target. Meanwhile, it effectively resolves the interdependencies among multi-scale encoder features in channel and spatial dimensions, which bridges the semantic gap between encoder and decoder features. Finally, a detail enhancement decoder module is designed to restore the resolution of the extracted feature maps, with the aim of addressing the issue of image detail information loss during the upsampling phase. This module builds upon the principles of DCD and incorporates a cascade upsampling(CU)module. The CU is specifically engineered to capture richer semantic information, retain feature details and semantic integrity, and ultimately ensure the high accuracy and delicate precision of the segmentation results. Our network achieves a highly specialized and nuanced segmentation of remote sensing building images by integrating these sophisticated components. Result Experimental results demonstrate the exceptional robustness of the GPDEA-UNet network introduced in this study across various datasets. Specifically, on the WHU Aerial Imagery Dataset(WHU), the network achieves an intersection over union(IoU)of 91. 60%, precision of 95. 36%, recall of 95. 89%, and an F1-score of 95. 62%. Similarly, on the Massachusetts Building Dataset, the network attains an IoU of 73. 51%, precision of 79. 44%, recall of 86. 81%, and an F1-score of 82. 53%. When compared with other state-of-the-art networks, the quantitative indicators reveal that the GPDEA-UNet network attains optimal performance on the WHU dataset and either optimal or near-optimal performance on the Massachusetts Building Dataset. Furthermore, qualitative analysis demonstrates that the proposed network achieves superior segmentation results on the WHU and Massachusetts Building Datasets. The network maintains high-quality segmentation even for remote sensing images with inferior imaging quality, such as those with low resolution, noise, or occlusion. Conclusion An asymmetric remote sensing building segmentation network with global perception and detail enhancement is proposed by combining a selective state space module and a multi-scale dual cross-attention mechanism. Experiments on two remote sensing datasets show that the proposed network can effectively improve the accuracy and visualization effects of remote sensing building segmentation. Furthermore, the network exhibits remarkable robustness and versatility. The high precision and recall rates achieved in our experiments highlight its capability to excel not only in high-quality remote sensing building segmentation but also in challenging scenarios. This study shows that the proposed network has excellent universality and application potential in remote sensing image segmentation and provides a new research idea and method for the research and application of remote sensing image processing.

Translated title of the contributionGlobal perception and detail enhancement network for building segmentation in remote sensing images
Original languageChinese (Traditional)
Pages (from-to)2866-2883
Number of pages18
JournalJournal of Image and Graphics
Volume30
Issue number8
DOIs
StatePublished - 2025

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 11 - Sustainable Cities and Communities
    SDG 11 Sustainable Cities and Communities

Fingerprint

Dive into the research topics of 'Global perception and detail enhancement network for building segmentation in remote sensing images'. Together they form a unique fingerprint.

Cite this