跳到主要导航 跳到搜索 跳到主要内容

MG-TVMF: Multi-grained text-video matching and fusing for weakly supervised video anomaly detection

  • Xi'an Jiaotong University

科研成果: 期刊稿件文章同行评审

2 引用 (Scopus)

摘要

Weakly supervised video anomaly detection (WS-VAD) often suffers from false alarms and incomplete localization due to the lack of precise temporal annotations. To address these limitations, we propose a novel method, multi-grained text-video matching and fusing (MG-TVMF), which leverages semantic cues from anomaly category text labels to enhance both the accuracy and completeness of anomaly localization. MG-TVMF integrates two complementary branches: the MG-TVM branch improves localization accuracy through a hierarchical structure comprising a coarse-grained classification module and two fine-grained matching modules, including a video-text matching (VTM) module for global semantic alignment and a segment-text matching (STM) module for local video (i.e. segment) text alignment via optimal transport algorithm. Meanwhile, the MG-TVF branch enhances localization completeness by prepending a global video-level text prompt to each segment-level caption for multi-grained textual fusion, and reconstructing the masked anomaly-related caption of the top-scoring segment using video segment features and anomaly scores. Extensive experiments on the UCF-Crime and XD-Violence datasets demonstrate the effectiveness of the proposed VTM and STM modules as well as the MG-TVF branch, and the proposed MG-TVMF method achieves state-of-the-art performance on UCF-Crime, XD-Violence, and ShanghaiTech datasets.

源语言英语
期刊论文编号113201
期刊Pattern Recognition
176
DOI
出版状态已出版 - 8月 2026

联合国可持续发展目标

此成果有助于实现下列可持续发展目标:

  1. 可持续发展目标 16 - 和平、正义和强大机构
    可持续发展目标 16 和平、正义和强大机构

学术指纹

探究 'MG-TVMF: Multi-grained text-video matching and fusing for weakly supervised video anomaly detection' 的科研主题。它们共同构成独一无二的学术指纹。

引用此