TY - GEN
T1 - Research on Small Target Detection Algorithm for Outdoor Complex Environment Based on STB-YOLOv8
AU - Liu, Ruosong
AU - Yang, Qingyu
AU - Li, Donghe
AU - Song, Pengtao
N1 - Publisher Copyright:
© 2025 Technical Committee on Control Theory, Chinese Association of Automation.
PY - 2025
Y1 - 2025
N2 - The task of object detection in complex outdoor scenes faces multiple challenges: differences in lighting conditions, dynamic changes in the object background, and incomplete targets due to occlusion. Existing research often focuses on the YOLO algorithm and optimizes it within the field of convolutional neural networks. However, the research on applying the Transformer algorithm to the field of image processing has a theoretical basis and good prospects, but still lacks extensive research and experiments. This paper proposes an improved YOLOv8n architecture, which replaces the C2f module located in the deep part of Backbone with Swin Transformer Block (STB) to take advantage of Transformer's global feature extraction capability. In addition, a small target detection head is added to capture the rich location information in the shallow layer of the network to improve the detection performance of small target objects. Simulation results show that the improved algorithm can effectively improve the detection effect, and the recall rate, F1-score, mAP50 and mAP50-95 are better than the original YOLOv8n model, increasing by 1.961%, 1.375%, 0.676%, 3.82% respectively.
AB - The task of object detection in complex outdoor scenes faces multiple challenges: differences in lighting conditions, dynamic changes in the object background, and incomplete targets due to occlusion. Existing research often focuses on the YOLO algorithm and optimizes it within the field of convolutional neural networks. However, the research on applying the Transformer algorithm to the field of image processing has a theoretical basis and good prospects, but still lacks extensive research and experiments. This paper proposes an improved YOLOv8n architecture, which replaces the C2f module located in the deep part of Backbone with Swin Transformer Block (STB) to take advantage of Transformer's global feature extraction capability. In addition, a small target detection head is added to capture the rich location information in the shallow layer of the network to improve the detection performance of small target objects. Simulation results show that the improved algorithm can effectively improve the detection effect, and the recall rate, F1-score, mAP50 and mAP50-95 are better than the original YOLOv8n model, increasing by 1.961%, 1.375%, 0.676%, 3.82% respectively.
KW - Global feature extraction
KW - Object detection
KW - Outdoor complex scenes
KW - Transformer
KW - YOLO
UR - https://www.scopus.com/pages/publications/105020309601
U2 - 10.23919/CCC64809.2025.11178405
DO - 10.23919/CCC64809.2025.11178405
M3 - 会议稿件
AN - SCOPUS:105020309601
T3 - Chinese Control Conference, CCC
SP - 7449
EP - 7454
BT - Proceedings of the 44th Chinese Control Conference, CCC 2025
A2 - Sun, Jian
A2 - Yin, Hongpeng
PB - IEEE Computer Society
T2 - 44th Chinese Control Conference, CCC 2025
Y2 - 28 July 2025 through 30 July 2025
ER -