Abstract
Accurate pedestrian crossing intention prediction is critical for ensuring safety and efficiency in autonomous driving. Existing approaches primarily rely on single-modality analysis (e.g., posture or trajectory) and struggle in modeling dynamic interactions between pedestrians, vehicles, and environmental contexts. To address these limitations, we propose SafeCrossNet, a novel multi-modal fusion framework that integrates social-aware with hierarchical spatio-temporal feature learning. Specifically, we design a scene interaction encoding module that employs a hybrid architecture combining Convolutional Neural Networks (CNN) and Gated Recurrent Units (GRU), which is enhanced with attention mechanisms to better capture the interactive relationships between pedestrians and environmental objects. Additionally, we introduce dynamic encoding module, for vehicle-side non-visual features, we adopt stacked multi-layer GRU augmented with attention mechanisms to achieve refined feature learning. Finally, we design a Hierarchical Spatio-temporal Synergistic Fusion (HSSF) strategy. This strategy facilitates adaptive reasoning across multi-modal features, significantly enhancing pedestrian crossing intention prediction accuracy. Extensive experiments conducted on widely used benchmarks, including PIE and JAAD datasets, demonstrate that our method achieves state-of-the-art performance, attaining accuracies of 91% on the PIE dataset and 90% on the JAAD dataset.
| Original language | English |
|---|---|
| Article number | 103609 |
| Journal | Information Fusion |
| Volume | 126 |
| DOIs | |
| State | Published - Feb 2026 |
| Externally published | Yes |
Keywords
- Attention mechanism
- Autonomous driving
- Gated Recurrent Unit (GRU)
- Multi-modal fusion
- Pedestrian crossing intention prediction
Fingerprint
Dive into the research topics of 'SafeCrossNet: Multi-modal fusion with social-aware for pedestrian crossing intention prediction'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver