Abstract
Egocentricly comprehending the causes and effects of car accidents is crucial for the safety of self-driving cars, and synthesizing causal-entity reflected accident videos can facilitate the capability test to respond to unaffordable accidents in reality. However, incorporating causal relations as seen in real-world videos into synthetic videos remains challenging. This work argues that precisely identifying the accident participants and capturing their related behaviors are of critical importance. In this regard, we propose a novel diffusion model Causal-VidSyn for synthesizing egocentric traffic accident videos. To enable causal entity grounding in video diffusion, Causal-VidSyn leverages the cause descriptions and driver fixations to identify the accident participants and behaviors, facilitated by accident reason answering and gaze-conditioned selection modules. To support CausalVidSyn, we further construct Drive-Gaze, the largest driver gaze dataset (with 1.54M frames of fixations) in driving accident scenarios. Extensive experiments show that CausalVidSyn surpasses state-of-the-art video diffusion models in terms of frame quality and causal sensitivity in various tasks, including accident video editing, normal-to-accident video diffusion, and text-to-video generation.
| Original language | English |
|---|---|
| Title of host publication | Proceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025 |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| Pages | 11208-11218 |
| Number of pages | 11 |
| ISBN (Electronic) | 9798331587758 |
| DOIs | |
| State | Published - 2025 |
| Event | 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025 - Honolulu, United States Duration: 19 Oct 2025 → 23 Oct 2025 |
Publication series
| Name | Proceedings of the IEEE International Conference on Computer Vision |
|---|---|
| ISSN (Print) | 1550-5499 |
| ISSN (Electronic) | 2380-7504 |
Conference
| Conference | 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025 |
|---|---|
| Country/Territory | United States |
| City | Honolulu |
| Period | 19/10/25 → 23/10/25 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- accident video synthesis
- causal reasoning
- driver attention
- video diffusion models
Fingerprint
Dive into the research topics of 'Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver