TY - GEN
T1 - Difference-Guided Modality Fusion Network for Multimodal Object Detection
AU - Li, Linxuan
AU - Liu, Meiqin
AU - Lan, Jian
AU - Dong, Shanling
AU - Liu, Zhunga
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - In recent years, visible-infrared object detection has achieved significant progress. However, most existing methods primarily emphasize the shared features between the two modalities while overlooking their feature differences. To address this limitation, we propose the Difference-Guided Modality Fusion Network, which can effectively improve the fusion and detection performance of modalities. Specifically, we propose a cross-modal data augmentation strategy to overcome the limitations of single-modality reliance by exchanging the partial modal information. To further capture and analyze feature differences between modalities, we introduce a differential attention fusion approach that models a difference matrix across modal channels, thereby quantifying and strengthening the salient features of the two modalities. Additionally, we develop a modality-aware dynamic learning mechanism that employs a loss function that can simultaneously focus on the differences and common parts of the modalities, guiding the model to adaptively learn features between the modalities. Experimental results on FLIR, LLVIP and M3FD datasets demonstrate the effectiveness of the proposed method, with mAP reaching 42.3%, 67.5% and 59.0% respectively.
AB - In recent years, visible-infrared object detection has achieved significant progress. However, most existing methods primarily emphasize the shared features between the two modalities while overlooking their feature differences. To address this limitation, we propose the Difference-Guided Modality Fusion Network, which can effectively improve the fusion and detection performance of modalities. Specifically, we propose a cross-modal data augmentation strategy to overcome the limitations of single-modality reliance by exchanging the partial modal information. To further capture and analyze feature differences between modalities, we introduce a differential attention fusion approach that models a difference matrix across modal channels, thereby quantifying and strengthening the salient features of the two modalities. Additionally, we develop a modality-aware dynamic learning mechanism that employs a loss function that can simultaneously focus on the differences and common parts of the modalities, guiding the model to adaptively learn features between the modalities. Experimental results on FLIR, LLVIP and M3FD datasets demonstrate the effectiveness of the proposed method, with mAP reaching 42.3%, 67.5% and 59.0% respectively.
UR - https://www.scopus.com/pages/publications/105033152512
U2 - 10.1109/SMC58881.2025.11343398
DO - 10.1109/SMC58881.2025.11343398
M3 - 会议稿件
AN - SCOPUS:105033152512
T3 - Conference Proceedings - IEEE International Conference on Systems, Man and Cybernetics
SP - 716
EP - 721
BT - 2025 IEEE International Conference on Systems, Man, and Cybernetics
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 IEEE International Conference on Systems, Man, and Cybernetics, SMC 2025
Y2 - 5 October 2025 through 8 October 2025
ER -