TY - JOUR
T1 - Self-Bidirectional Decoupled Distillation for Time Series Classification
AU - Xiao, Zhiwen
AU - Xing, Huanlai
AU - Qu, Rong
AU - Li, Hui
AU - Feng, Li
AU - Zhao, Bowen
AU - Yang, Jiayi
N1 - Publisher Copyright:
© 2020 IEEE.
PY - 2024
Y1 - 2024
N2 - Over the years, many deep learning algorithms have been developed for time series classification (TSC). A learning model's performance usually depends on the quality of the semantic information extracted from lower and higher levels within the representation hierarchy. Efficiently promoting mutual learning between higher and lower levels is vital to enhance the model's performance during model learning. To this end, we propose a self-bidirectional decoupled distillation (self-BiDecKD) method for TSC. Unlike most self-distillation algorithms that usually transfer the target-class knowledge from higher to lower levels, self-BiDecKD encourages the output of the output layer and the output of each lower level block to form a bidirectional decoupled knowledge distillation (KD) pair. The bidirectional decoupled KD promotes mutual learning between lower and higher level semantic information and extracts the knowledge hidden in the target and nontarget classes, helping self-BiDecKD capture rich representations from the data. Experimental results show that compared with a number of self-distillation algorithms, self-BiDecKD wins 35 out of 85 University of California, Riverside (UCR) 2018 datasets and achieves the smallest AVeraGe (AVG)_rank score, namely 3.2882. In particular, compared with a nonself-distillation baseline, self-BiDecKD results in 58/8/19 regarding 'win'/'tie'/'lose.'
AB - Over the years, many deep learning algorithms have been developed for time series classification (TSC). A learning model's performance usually depends on the quality of the semantic information extracted from lower and higher levels within the representation hierarchy. Efficiently promoting mutual learning between higher and lower levels is vital to enhance the model's performance during model learning. To this end, we propose a self-bidirectional decoupled distillation (self-BiDecKD) method for TSC. Unlike most self-distillation algorithms that usually transfer the target-class knowledge from higher to lower levels, self-BiDecKD encourages the output of the output layer and the output of each lower level block to form a bidirectional decoupled knowledge distillation (KD) pair. The bidirectional decoupled KD promotes mutual learning between lower and higher level semantic information and extracts the knowledge hidden in the target and nontarget classes, helping self-BiDecKD capture rich representations from the data. Experimental results show that compared with a number of self-distillation algorithms, self-BiDecKD wins 35 out of 85 University of California, Riverside (UCR) 2018 datasets and achieves the smallest AVeraGe (AVG)_rank score, namely 3.2882. In particular, compared with a nonself-distillation baseline, self-BiDecKD results in 58/8/19 regarding 'win'/'tie'/'lose.'
KW - Convolutional neural network
KW - data mining
KW - deep learning
KW - knowledge distillation (KD)
KW - time series classification (TSE)
UR - https://www.scopus.com/pages/publications/85184342102
U2 - 10.1109/TAI.2024.3360180
DO - 10.1109/TAI.2024.3360180
M3 - 文章
AN - SCOPUS:85184342102
SN - 2691-4581
VL - 5
SP - 4101
EP - 4110
JO - IEEE Transactions on Artificial Intelligence
JF - IEEE Transactions on Artificial Intelligence
IS - 8
ER -