Skip to main navigation Skip to search Skip to main content

MG-TVMF: Multi-grained text-video matching and fusing for weakly supervised video anomaly detection

  • Xi'an Jiaotong University

Research output: Contribution to journalArticlepeer-review

2 Scopus citations

Abstract

Weakly supervised video anomaly detection (WS-VAD) often suffers from false alarms and incomplete localization due to the lack of precise temporal annotations. To address these limitations, we propose a novel method, multi-grained text-video matching and fusing (MG-TVMF), which leverages semantic cues from anomaly category text labels to enhance both the accuracy and completeness of anomaly localization. MG-TVMF integrates two complementary branches: the MG-TVM branch improves localization accuracy through a hierarchical structure comprising a coarse-grained classification module and two fine-grained matching modules, including a video-text matching (VTM) module for global semantic alignment and a segment-text matching (STM) module for local video (i.e. segment) text alignment via optimal transport algorithm. Meanwhile, the MG-TVF branch enhances localization completeness by prepending a global video-level text prompt to each segment-level caption for multi-grained textual fusion, and reconstructing the masked anomaly-related caption of the top-scoring segment using video segment features and anomaly scores. Extensive experiments on the UCF-Crime and XD-Violence datasets demonstrate the effectiveness of the proposed VTM and STM modules as well as the MG-TVF branch, and the proposed MG-TVMF method achieves state-of-the-art performance on UCF-Crime, XD-Violence, and ShanghaiTech datasets.

Original languageEnglish
Article number113201
JournalPattern Recognition
Volume176
DOIs
StatePublished - Aug 2026

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 16 - Peace, Justice and Strong Institutions
    SDG 16 Peace, Justice and Strong Institutions

Keywords

  • Multi-grained text-video fusing
  • Multi-grained text-video matching
  • Optimal transport
  • Weakly-supervised anomaly detection

Fingerprint

Dive into the research topics of 'MG-TVMF: Multi-grained text-video matching and fusing for weakly supervised video anomaly detection'. Together they form a unique fingerprint.

Cite this