Skip to main navigation Skip to search Skip to main content

SFNet: Sparse Fusion Network for Object Detection of Cross-Modal Images

  • Haidong Xiao
  • , Zhigang Ren
  • , Ziyu Li
  • , Zhuoxun Zeng
  • , Shuangping Yang
  • , Zimu Teng
  • , Yijie Wang
  • , Shengze Cai
  • , Chao Xu
  • , Zongze Wu
  • Guangdong University of Technology
  • Zhejiang University
  • Sichuan Aerospace System Engineering Institute
  • Shenzhen University

Research output: Contribution to journalArticlepeer-review

Abstract

Effective aggregation of complementary information from visible and infrared modalities can significantly enhance the performance and robustness of multimodal object detection systems. However, existing methods often suffer from limitations, such as overly complex architectures or ineffective cross-modal interaction at the semantic level. Furthermore, some approaches integrate features extensively without adequate filtering mechanisms, potentially introducing interference from redundant information. To address these challenges, we propose SFNet, an end-to-end cross-modal object detection network based on a sparse fusion strategy. SFNet utilizes a weight-shared twinned backbone backbone to synchronously encode feature maps from both modalities. We introduce a novel Sparse Fusion Module (SFM) that operates at the semantic level to refine salient cross-modal features while simultaneously filtering redundant components during information interaction. Additionally, an Adaptive Proportional Modulation Module (APMM) is incorporated to dynamically adjust attention weights based on the characteristics of the fused multi-modal feature distribution. Extensive qualitative and quantitative experiments conducted on several benchmark datasets (M3FD, KAIST, FLIR) validate the effectiveness and superiority of SFNet. Our method achieves mean Average Precision (mAP) scores of 90.4% on theM3FD dataset, 78.3% on the KAIST dataset and 87. 0% on the FLIR dataset.

Original languageEnglish
JournalIEEE Sensors Journal
DOIs
StateAccepted/In press - 2026
Externally publishedYes

Keywords

  • cross-modal
  • feature fusion
  • feature screening
  • multispectral
  • object detection
  • weight modulation

Fingerprint

Dive into the research topics of 'SFNet: Sparse Fusion Network for Object Detection of Cross-Modal Images'. Together they form a unique fingerprint.

Cite this