跳到主要导航 跳到搜索 跳到主要内容

Memory-Efficient Batch Normalization by One-Pass Computation for On-Device Training

  • Xi'an Jiaotong University

科研成果: 期刊稿件文章同行评审

4 引用 (Scopus)

摘要

Batch normalization (BN) has become ubiquitous in modern deep learning architectures because of its remarkable improvement in deep neural network (DNN) training performance. However, the two-pass computation of statistical estimation and element-wise normalization in BN training requires two accesses to the input data, resulting in a huge increase in off-chip memory traffic during DNN training. In this brief, we propose a novel accelerator, named one-pass normalizer (OPN) to achieve memory-efficient BN for on-device training. Specifically, in terms of dataflow, we propose one-pass computation based on sampling-based range normalization and sparse data recovery techniques to reduce BN off-chip memory access. Regarding the OPN circuit, we propose channel-wise constant extraction to achieve a compact design. Experimental results show that the one-pass computation reduces off-chip memory access of BN by 2.0~3.8× compared with the previous state-of-the-art designs while maintaining training performance. Moreover, the channel-wise constant extraction saves the gate count and power consumption of OPN by 56% and 73%, respectively.

源语言英语
页(从-至)3186-3190
页数5
期刊IEEE Transactions on Circuits and Systems II: Express Briefs
71
6
DOI
出版状态已出版 - 1 6月 2024

学术指纹

探究 'Memory-Efficient Batch Normalization by One-Pass Computation for On-Device Training' 的科研主题。它们共同构成独一无二的学术指纹。

引用此