Skip to main navigation Skip to search Skip to main content

Caption Generation From Road Images for Traffic Scene Modeling

  • Yaochen Li
  • , Chuan Wu
  • , Ling Li
  • , Yuehu Liu
  • , Jihua Zhu
  • Xi'an Jiaotong University

Research output: Contribution to journalArticlepeer-review

17 Scopus citations

Abstract

In this traffic-scene-modeling study, we propose an image-captioning network which incorporates element attention into an encoder-decoder mechanism to generate more reasonable scene captions. A visual-relationship-detecting network is also developed to detect the relative positions of object pairs. Firstly, the traffic scene elements are detected and segmented according to their clustered locations. Then, the image-captioning network is applied to generate the corresponding description of each traffic scene element. The visual-relationship-detecting network is utilized to detect the position relations of all object pairs in the subregion. The static and dynamic traffic elements are appropriately selected and organized to construct a 3D model according to the captions and the position relations. The reconstructed 3D traffic scenes can be utilized for the offline test of unmanned vehicles. The evaluations and comparisons based on the TSD-max, KITTI and Microsoft's COCO datasets demonstrate the effectiveness of the proposed framework.

Original languageEnglish
Pages (from-to)7805-7816
Number of pages12
JournalIEEE Transactions on Intelligent Transportation Systems
Volume23
Issue number7
DOIs
StatePublished - 1 Jul 2022

Keywords

  • Geometric analysis
  • image captioning
  • road scene layout
  • traffic scene construction
  • visual relationship detection

Fingerprint

Dive into the research topics of 'Caption Generation From Road Images for Traffic Scene Modeling'. Together they form a unique fingerprint.

Cite this