TY - GEN
T1 - BBP-DNS
T2 - 2025 International Symposium of Electronics Design Automation, ISEDA 2025
AU - Yuan, Shuai
AU - Niu, Dan
AU - Zhao, Huatao
AU - Liu, Shiyuan
AU - Jin, Zhou
AU - Sun, Changyin
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Transformer-based GPT models in artificial intelligence have exhibited remarkable performance advantages across generative tasks. However, edge-side deployment of GPT models faces significant challenges due to the limited memory and computational resources of edge devices. The lack of specialized compiler toolchains often necessitates manual compilation, resulting in complicated deployment processes. To address these challenges, we propose a Batch-Block Parallelization and Dual-NoC Scheduling algorithm(BBP-DNS) for edge devices, which automates tensor partitioning and batch scheduling to enhance data locality and reduce overheads in data storage, movement, and kernel scheduling, thereby improving inference speed. To mitigate hardware constraints in accelerator memory, bandwidth, and compute resources, we extend tensor parallelism with a fine-grained batch-level scheduling strategy. Additionally, we introduce a novel operator mapping methodology that automates accelerator data management, addressing the inefficiencies of traditional manual compilation workflow. Experimental results demonstrate BBP-DNS's superior performance, achieving a 30x performance improvement of a GPT single-layer block on simulator of that evaluated by TVM and PyTorch on CPUs.
AB - Transformer-based GPT models in artificial intelligence have exhibited remarkable performance advantages across generative tasks. However, edge-side deployment of GPT models faces significant challenges due to the limited memory and computational resources of edge devices. The lack of specialized compiler toolchains often necessitates manual compilation, resulting in complicated deployment processes. To address these challenges, we propose a Batch-Block Parallelization and Dual-NoC Scheduling algorithm(BBP-DNS) for edge devices, which automates tensor partitioning and batch scheduling to enhance data locality and reduce overheads in data storage, movement, and kernel scheduling, thereby improving inference speed. To mitigate hardware constraints in accelerator memory, bandwidth, and compute resources, we extend tensor parallelism with a fine-grained batch-level scheduling strategy. Additionally, we introduce a novel operator mapping methodology that automates accelerator data management, addressing the inefficiencies of traditional manual compilation workflow. Experimental results demonstrate BBP-DNS's superior performance, achieving a 30x performance improvement of a GPT single-layer block on simulator of that evaluated by TVM and PyTorch on CPUs.
KW - In-Memory Computing
KW - Model Deployment
KW - Network-on-Chip
KW - Neural Network
KW - Tensor Parallelism
UR - https://www.scopus.com/pages/publications/105014241988
U2 - 10.1109/ISEDA65950.2025.11100292
DO - 10.1109/ISEDA65950.2025.11100292
M3 - 会议稿件
AN - SCOPUS:105014241988
T3 - 2025 International Symposium of Electronics Design Automation, ISEDA 2025
SP - 114
EP - 119
BT - 2025 International Symposium of Electronics Design Automation, ISEDA 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 9 May 2025 through 12 May 2025
ER -