Skip to main navigation Skip to search Skip to main content

SafeCrossNet: Multi-modal fusion with social-aware for pedestrian crossing intention prediction

  • Quancheng Du
  • , Lele Xu
  • , Qiong Wu
  • , Huansheng Ning
  • , Xiao Wang
  • , Liang Lin
  • , Changyin Sun
  • University of Science and Technology Beijing
  • Anhui University
  • Anhui Jianghuai Automobile Co., Ltd.
  • Anhui Province Engineering Center of Unmanned Systems and Intelligent Technology
  • Sun Yat-Sen University
  • Key Lab of the Ministry of Education for Process Control and Efficiency Egineering

Research output: Contribution to journalArticlepeer-review

6 Scopus citations

Abstract

Accurate pedestrian crossing intention prediction is critical for ensuring safety and efficiency in autonomous driving. Existing approaches primarily rely on single-modality analysis (e.g., posture or trajectory) and struggle in modeling dynamic interactions between pedestrians, vehicles, and environmental contexts. To address these limitations, we propose SafeCrossNet, a novel multi-modal fusion framework that integrates social-aware with hierarchical spatio-temporal feature learning. Specifically, we design a scene interaction encoding module that employs a hybrid architecture combining Convolutional Neural Networks (CNN) and Gated Recurrent Units (GRU), which is enhanced with attention mechanisms to better capture the interactive relationships between pedestrians and environmental objects. Additionally, we introduce dynamic encoding module, for vehicle-side non-visual features, we adopt stacked multi-layer GRU augmented with attention mechanisms to achieve refined feature learning. Finally, we design a Hierarchical Spatio-temporal Synergistic Fusion (HSSF) strategy. This strategy facilitates adaptive reasoning across multi-modal features, significantly enhancing pedestrian crossing intention prediction accuracy. Extensive experiments conducted on widely used benchmarks, including PIE and JAAD datasets, demonstrate that our method achieves state-of-the-art performance, attaining accuracies of 91% on the PIE dataset and 90% on the JAAD dataset.

Original languageEnglish
Article number103609
JournalInformation Fusion
Volume126
DOIs
StatePublished - Feb 2026
Externally publishedYes

Keywords

  • Attention mechanism
  • Autonomous driving
  • Gated Recurrent Unit (GRU)
  • Multi-modal fusion
  • Pedestrian crossing intention prediction

Fingerprint

Dive into the research topics of 'SafeCrossNet: Multi-modal fusion with social-aware for pedestrian crossing intention prediction'. Together they form a unique fingerprint.

Cite this