跳到主要导航 跳到搜索 跳到主要内容

面向多功能张量加速器的细粒度结构化稀疏设计

  • Huazheng Zhao
  • , Shanmin Pang
  • , Yinghai Zhao
  • , Gaohui Hua
  • , Chenyang Li
  • , Zhansheng Duan
  • , Kuizhi Mei
  • Xi'an Jiaotong University
  • Beijing Huahang Institute of Radio Measurement

科研成果: 期刊稿件文章同行评审

摘要

In order to address the compatibility issue between model compression algorithms and the versatile tensor accelerator (VTA) , an adaptive fine-grained structured sparse design tailored for this accelerator is proposed by enhancing the classical YOLObile block-wise pruning method and evaluates its performance. In light of the multi-dimensional loop unfolding characteristics of VTA, the model's weight tensors are divided into 32X32 blocks. This approach integrates temporal distillation and spatial distillation to align multidimensional features. Through a single-stage iterative training method, the calculation process of the original ADMM algorithm is refined to improve model deployment accuracy while reducing training costs. An adaptive layer pruning rate module is introduced to dynamically allocate the total pruning rate, facilitating end-to-end automated pruning. The experimental results demonstrate that this improved method effectively reduces floating-point computations by approximately 2.4% and enhances the accuracy of compressed models across various tasks such as image classification and object detection, with a maximum growth percentage of 2. 6%. This method offers an efficient and lightweight software solution for the sparse deployment of deep learning models on VTAs.

投稿的翻译标题Fine-Grained Structured Sparse Design for Versatile Tensor Accelerator
源语言繁体中文
页(从-至)176-184
页数9
期刊Hsi-An Chiao Tung Ta Hsueh/Journal of Xi'an Jiaotong University
58
11
DOI
出版状态已出版 - 11月 2024

关键词

  • deep learning
  • model deployment
  • model sparsity
  • neural network compression
  • versatile tensor accelerator

学术指纹

探究 '面向多功能张量加速器的细粒度结构化稀疏设计' 的科研主题。它们共同构成独一无二的学术指纹。

引用此