Skip to main navigation Skip to search Skip to main content

PR-DETR: Injecting position and relation prior for dense video captioning

  • State Key Laboratory of Human–Machine Hybrid Augmented Intelligence
  • Xi'an Jiaotong University

Research output: Contribution to journalArticlepeer-review

Abstract

Dense video captioning is a challenging task that aims to localize and caption multiple events in an untrimmed video. Recent studies mainly follow the transformer-based architecture to jointly perform the two sub-tasks, i.e., event localization and caption generation, in an end-to-end manner. Based on the general philosophy of detection transformer, these methods implicitly learn the event locations and event semantics, which requires a large amount of training data and limits the model's performance in practice. In this paper, we propose a novel dense video captioning framework, named PR-DETR, which injects the explicit position and relation prior into the detection transformer to improve the localization accuracy and caption quality, simultaneously. On the one hand, we first generate a set of position-anchored queries to provide the scene-specific position and semantics of potential events as position prior, which serves as the initial event search regions to eliminate the implausible event proposals. On the other hand, we further design an event relation encoder to explicitly calculate the relationships between event boundaries and encode them into the relation mask as relation prior to improve the coherence of the captions. Extensive ablation studies are conducted to verify the effectiveness of the position and relation prior. Experimental results also show the competitive performance of our method on ActivityNet Captions and YouCook2 datasets.

Original languageEnglish
Article number113872
JournalPattern Recognition
Volume179
DOIs
StatePublished - Nov 2026

Keywords

  • Dense video captioning
  • DETR
  • Prior information

Fingerprint

Dive into the research topics of 'PR-DETR: Injecting position and relation prior for dense video captioning'. Together they form a unique fingerprint.

Cite this