TY - GEN
T1 - Explore a Novel Knowledge Distillation Framework for Network Learning and Low-Bit Quantization
AU - Si, Liang
AU - Li, Yuhai
AU - Zhou, Hengyi
AU - Liang, Jiahua
AU - Liu, Longjun
N1 - Publisher Copyright:
© 2021 IEEE
PY - 2021
Y1 - 2021
N2 - Knowledge distillation is a kind of model compression methods. It improves the performance of “student” networks by transferring the knowledge of “teacher” to “student”. However, due to the huge gap between “teacher” and “student” in knowledge representation capabilities, directly minimizing the difference between them to transfer the knowledge will lead to the convergence issue. To this end, this paper proposes a novel framework of knowledge distillation. By sharing the weights of “teacher” and “student” in fully connected layers, we reduce the challenge of knowledge transfer. Furthermore, we propose a novel training strategy to improve the performance of low-bit quantization networks based on our distillation framework called distillation for low-bit quantization (DLBQ). The experimental results show that our methods can achieve significant improvement across different tasks. For instance, ResNet-20 gains 1.81% improvement on Cifar10 dataset. ResNet-56 shows 3.36% improvement on Cifar100 dataset and even exhibits 1.72% performance improvement than”teacher” network. Additionally, as for the improvements of quantization performance, ResNet-20 gain 1.25% improvement with ternary weights on Cifar10, and ResNet-110 manifects 3.22% improvement with binary weights on Cifar100 dataset.
AB - Knowledge distillation is a kind of model compression methods. It improves the performance of “student” networks by transferring the knowledge of “teacher” to “student”. However, due to the huge gap between “teacher” and “student” in knowledge representation capabilities, directly minimizing the difference between them to transfer the knowledge will lead to the convergence issue. To this end, this paper proposes a novel framework of knowledge distillation. By sharing the weights of “teacher” and “student” in fully connected layers, we reduce the challenge of knowledge transfer. Furthermore, we propose a novel training strategy to improve the performance of low-bit quantization networks based on our distillation framework called distillation for low-bit quantization (DLBQ). The experimental results show that our methods can achieve significant improvement across different tasks. For instance, ResNet-20 gains 1.81% improvement on Cifar10 dataset. ResNet-56 shows 3.36% improvement on Cifar100 dataset and even exhibits 1.72% performance improvement than”teacher” network. Additionally, as for the improvements of quantization performance, ResNet-20 gain 1.25% improvement with ternary weights on Cifar10, and ResNet-110 manifects 3.22% improvement with binary weights on Cifar100 dataset.
KW - Deep Neural Networks
KW - Knowledge Distillation
KW - Low-Bit Quantization
UR - https://www.scopus.com/pages/publications/85128046194
U2 - 10.1109/CAC53003.2021.9728523
DO - 10.1109/CAC53003.2021.9728523
M3 - 会议稿件
AN - SCOPUS:85128046194
T3 - Proceeding - 2021 China Automation Congress, CAC 2021
SP - 3002
EP - 3007
BT - Proceeding - 2021 China Automation Congress, CAC 2021
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2021 China Automation Congress, CAC 2021
Y2 - 22 October 2021 through 24 October 2021
ER -