Skip to main navigation Skip to search Skip to main content

Generative Model-Based Feature Knowledge Distillation for Action Recognition

  • Xi'an Jiaotong University

Research output: Contribution to journalConference articlepeer-review

11 Scopus citations

Abstract

Knowledge distillation (KD), a technique widely employed in computer vision, has emerged as a de facto standard for improving the performance of small neural networks. However, prevailing KD-based approaches in video tasks primarily focus on designing loss functions and fusing cross-modal information. This overlooks the spatial-temporal feature semantics, resulting in limited advancements in model compression. Addressing this gap, our paper introduces an innovative knowledge distillation framework, with the generative model for training a lightweight student model. In particular, the framework is organized into two steps: the initial phase is Feature Representation, wherein a generative model-based attention module is trained to represent feature semantics; Subsequently, the Generative-based Feature Distillation phase encompasses both Generative Distillation and Attention Distillation, with the objective of transferring attention-based feature semantics with the generative model. The efficacy of our approach is demonstrated through comprehensive experiments on diverse popular datasets, proving considerable enhancements in video action recognition task. Moreover, the effectiveness of our proposed framework is validated in the context of more intricate video action detection task. Our code is available at https://github.com/aaai24/Generative-based-KD.

Original languageEnglish
Pages (from-to)15474-15482
Number of pages9
JournalProceedings of the AAAI Conference on Artificial Intelligence
Volume38
Issue number14
DOIs
StatePublished - 25 Mar 2024
Event38th AAAI Conference on Artificial Intelligence, AAAI 2024 - Vancouver, Canada
Duration: 20 Feb 202427 Feb 2024

Fingerprint

Dive into the research topics of 'Generative Model-Based Feature Knowledge Distillation for Action Recognition'. Together they form a unique fingerprint.

Cite this