跳到主要导航 跳到搜索 跳到主要内容

Learning Visual-Audio Dissonance for Moment Retrieval and Highlight Detection

  • Xi'an Jiaotong University

科研成果: 期刊稿件文章同行评审

1 引用 (Scopus)

摘要

Language-guided video moment retrieval and highlight detection have made impressive advancements recently. Most existing methods primarily focus on visual cues in videos. In reality, audio cues in videos also play an indispensable role. In this paper, we highlight the importance of language-guided visual-audio temporal semantic dissonances and propose a novel framework named VABooster. It aims to leverage audio information as boosting cues to jointly address moment retrieval and highlight detection. We propose a visual-audio synergy network to learn task-specific visual-audio features. Additionally, we introduce an omni-modality interaction component with the dual branch structure to capture visual-audio-text contexts for different tasks. A saliency token adapter is designed to regulate the learning process of global saliency tokens. VABooster surpasses the existing state-of-the-art methods with fewer parameters. Extensive experiments and in-depth ablation studies on QVHighlights, Charades-STA, TVSum, and Youtube Highlights datasets substantiate the effectiveness and robustness of the proposed framework.

源语言英语
期刊IEEE Transactions on Multimedia
DOI
出版状态已接受/待刊 - 2026

学术指纹

探究 'Learning Visual-Audio Dissonance for Moment Retrieval and Highlight Detection' 的科研主题。它们共同构成独一无二的指纹。

引用此