TY - JOUR
T1 - SafeCrossNet
T2 - Multi-modal fusion with social-aware for pedestrian crossing intention prediction
AU - Du, Quancheng
AU - Xu, Lele
AU - Wu, Qiong
AU - Ning, Huansheng
AU - Wang, Xiao
AU - Lin, Liang
AU - Sun, Changyin
N1 - Publisher Copyright:
© 2025 Elsevier B.V.
PY - 2026/2
Y1 - 2026/2
N2 - Accurate pedestrian crossing intention prediction is critical for ensuring safety and efficiency in autonomous driving. Existing approaches primarily rely on single-modality analysis (e.g., posture or trajectory) and struggle in modeling dynamic interactions between pedestrians, vehicles, and environmental contexts. To address these limitations, we propose SafeCrossNet, a novel multi-modal fusion framework that integrates social-aware with hierarchical spatio-temporal feature learning. Specifically, we design a scene interaction encoding module that employs a hybrid architecture combining Convolutional Neural Networks (CNN) and Gated Recurrent Units (GRU), which is enhanced with attention mechanisms to better capture the interactive relationships between pedestrians and environmental objects. Additionally, we introduce dynamic encoding module, for vehicle-side non-visual features, we adopt stacked multi-layer GRU augmented with attention mechanisms to achieve refined feature learning. Finally, we design a Hierarchical Spatio-temporal Synergistic Fusion (HSSF) strategy. This strategy facilitates adaptive reasoning across multi-modal features, significantly enhancing pedestrian crossing intention prediction accuracy. Extensive experiments conducted on widely used benchmarks, including PIE and JAAD datasets, demonstrate that our method achieves state-of-the-art performance, attaining accuracies of 91% on the PIE dataset and 90% on the JAAD dataset.
AB - Accurate pedestrian crossing intention prediction is critical for ensuring safety and efficiency in autonomous driving. Existing approaches primarily rely on single-modality analysis (e.g., posture or trajectory) and struggle in modeling dynamic interactions between pedestrians, vehicles, and environmental contexts. To address these limitations, we propose SafeCrossNet, a novel multi-modal fusion framework that integrates social-aware with hierarchical spatio-temporal feature learning. Specifically, we design a scene interaction encoding module that employs a hybrid architecture combining Convolutional Neural Networks (CNN) and Gated Recurrent Units (GRU), which is enhanced with attention mechanisms to better capture the interactive relationships between pedestrians and environmental objects. Additionally, we introduce dynamic encoding module, for vehicle-side non-visual features, we adopt stacked multi-layer GRU augmented with attention mechanisms to achieve refined feature learning. Finally, we design a Hierarchical Spatio-temporal Synergistic Fusion (HSSF) strategy. This strategy facilitates adaptive reasoning across multi-modal features, significantly enhancing pedestrian crossing intention prediction accuracy. Extensive experiments conducted on widely used benchmarks, including PIE and JAAD datasets, demonstrate that our method achieves state-of-the-art performance, attaining accuracies of 91% on the PIE dataset and 90% on the JAAD dataset.
KW - Attention mechanism
KW - Autonomous driving
KW - Gated Recurrent Unit (GRU)
KW - Multi-modal fusion
KW - Pedestrian crossing intention prediction
UR - https://www.scopus.com/pages/publications/105013635583
U2 - 10.1016/j.inffus.2025.103609
DO - 10.1016/j.inffus.2025.103609
M3 - 文章
AN - SCOPUS:105013635583
SN - 1566-2535
VL - 126
JO - Information Fusion
JF - Information Fusion
M1 - 103609
ER -