TY - JOUR
T1 - Learning Visual-Audio Dissonance for Moment Retrieval and Highlight Detection
AU - Yang, Jin
AU - Wei, Ping
AU - Zheng, Nanning
N1 - Publisher Copyright:
© 1999-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Language-guided video moment retrieval and highlight detection have made impressive advancements recently. Most existing methods primarily focus on visual cues in videos. In reality, audio cues in videos also play an indispensable role. In this paper, we highlight the importance of language-guided visual-audio temporal semantic dissonances and propose a novel framework named VABooster. It aims to leverage audio information as boosting cues to jointly address moment retrieval and highlight detection. We propose a visual-audio synergy network to learn task-specific visual-audio features. Additionally, we introduce an omni-modality interaction component with the dual branch structure to capture visual-audio-text contexts for different tasks. A saliency token adapter is designed to regulate the learning process of global saliency tokens. VABooster surpasses the existing state-of-the-art methods with fewer parameters. Extensive experiments and in-depth ablation studies on QVHighlights, Charades-STA, TVSum, and Youtube Highlights datasets substantiate the effectiveness and robustness of the proposed framework.
AB - Language-guided video moment retrieval and highlight detection have made impressive advancements recently. Most existing methods primarily focus on visual cues in videos. In reality, audio cues in videos also play an indispensable role. In this paper, we highlight the importance of language-guided visual-audio temporal semantic dissonances and propose a novel framework named VABooster. It aims to leverage audio information as boosting cues to jointly address moment retrieval and highlight detection. We propose a visual-audio synergy network to learn task-specific visual-audio features. Additionally, we introduce an omni-modality interaction component with the dual branch structure to capture visual-audio-text contexts for different tasks. A saliency token adapter is designed to regulate the learning process of global saliency tokens. VABooster surpasses the existing state-of-the-art methods with fewer parameters. Extensive experiments and in-depth ablation studies on QVHighlights, Charades-STA, TVSum, and Youtube Highlights datasets substantiate the effectiveness and robustness of the proposed framework.
KW - multi-modal learning
KW - video highlight detection
KW - Video moment retrieval
KW - visual-audio-language fusion
UR - https://www.scopus.com/pages/publications/105036718977
U2 - 10.1109/TMM.2026.3685800
DO - 10.1109/TMM.2026.3685800
M3 - 文章
AN - SCOPUS:105036718977
SN - 1520-9210
JO - IEEE Transactions on Multimedia
JF - IEEE Transactions on Multimedia
ER -