TY - GEN
T1 - Feature Fusion Network Based on Hybrid Attention for Semantic Segmentation
AU - Xie, Xinchen
AU - Li, Chen
AU - Tian, Lihua
N1 - Publisher Copyright:
© 2022 IEEE.
PY - 2022
Y1 - 2022
N2 - In the deep learning based real-Time image semantic segmentation task, there are high requirements for the inference speed of the network. Due to the small amounts of parameters of the lightweight backbones, the calculation speed is often faster, which meets the requirements of real-Time tasks. However, the ability of the lightweight networks to extract features is relatively weak, resulting in much worse segmentation accuracy than the large model. Therefore, how to make full use of the lightweight networks to extract more image information to achieve better segmentation performance has become a key problem. Here, we propose an efficient feature fusion network based on attention mechanism. First, the widely used MobileNetV2 is selected as the lightweight backbone network, and then spatial attention and channel attention are calculated for both high-resolution low-level features and low-resolution high-level features, thus the final feature map got a global receptive field. Besides, through the multi-levels supervised learning for each stage of the backbone, the multi-stage auxiliary loss function enables the network to be trained more effectively. Finally, on the cityscapes dataset, the our proposed network reached 74.12% mIoU, and the inference speed remained at 110 fps.
AB - In the deep learning based real-Time image semantic segmentation task, there are high requirements for the inference speed of the network. Due to the small amounts of parameters of the lightweight backbones, the calculation speed is often faster, which meets the requirements of real-Time tasks. However, the ability of the lightweight networks to extract features is relatively weak, resulting in much worse segmentation accuracy than the large model. Therefore, how to make full use of the lightweight networks to extract more image information to achieve better segmentation performance has become a key problem. Here, we propose an efficient feature fusion network based on attention mechanism. First, the widely used MobileNetV2 is selected as the lightweight backbone network, and then spatial attention and channel attention are calculated for both high-resolution low-level features and low-resolution high-level features, thus the final feature map got a global receptive field. Besides, through the multi-levels supervised learning for each stage of the backbone, the multi-stage auxiliary loss function enables the network to be trained more effectively. Finally, on the cityscapes dataset, the our proposed network reached 74.12% mIoU, and the inference speed remained at 110 fps.
KW - Attention Mechanism
KW - Lightweight model
KW - Semantic Segmentation
UR - https://www.scopus.com/pages/publications/85134879183
U2 - 10.1109/AIIoT54504.2022.9817347
DO - 10.1109/AIIoT54504.2022.9817347
M3 - 会议稿件
AN - SCOPUS:85134879183
T3 - 2022 IEEE World AI IoT Congress, AIIoT 2022
SP - 9
EP - 14
BT - 2022 IEEE World AI IoT Congress, AIIoT 2022
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2022 IEEE World AI IoT Congress, AIIoT 2022
Y2 - 6 June 2022 through 9 June 2022
ER -