Skip to main navigation Skip to search Skip to main content

BBP-DNS: Batch-Block Parallelism and Dual-NoC Scheduling for Accelerating GPT on Edge Devices

  • Shuai Yuan
  • , Dan Niu
  • , Huatao Zhao
  • , Shiyuan Liu
  • , Zhou Jin
  • , Changyin Sun
  • Southeast University, Nanjing
  • NoC Team Beijing Institute of Open Source Chip
  • China University of Petroleum - Beijing
  • Anhui University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Transformer-based GPT models in artificial intelligence have exhibited remarkable performance advantages across generative tasks. However, edge-side deployment of GPT models faces significant challenges due to the limited memory and computational resources of edge devices. The lack of specialized compiler toolchains often necessitates manual compilation, resulting in complicated deployment processes. To address these challenges, we propose a Batch-Block Parallelization and Dual-NoC Scheduling algorithm(BBP-DNS) for edge devices, which automates tensor partitioning and batch scheduling to enhance data locality and reduce overheads in data storage, movement, and kernel scheduling, thereby improving inference speed. To mitigate hardware constraints in accelerator memory, bandwidth, and compute resources, we extend tensor parallelism with a fine-grained batch-level scheduling strategy. Additionally, we introduce a novel operator mapping methodology that automates accelerator data management, addressing the inefficiencies of traditional manual compilation workflow. Experimental results demonstrate BBP-DNS's superior performance, achieving a 30x performance improvement of a GPT single-layer block on simulator of that evaluated by TVM and PyTorch on CPUs.

Original languageEnglish
Title of host publication2025 International Symposium of Electronics Design Automation, ISEDA 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages114-119
Number of pages6
ISBN (Electronic)9798331536961
DOIs
StatePublished - 2025
Externally publishedYes
Event2025 International Symposium of Electronics Design Automation, ISEDA 2025 - Hong Kong, China
Duration: 9 May 202512 May 2025

Publication series

Name2025 International Symposium of Electronics Design Automation, ISEDA 2025

Conference

Conference2025 International Symposium of Electronics Design Automation, ISEDA 2025
Country/TerritoryChina
CityHong Kong
Period9/05/2512/05/25

Keywords

  • In-Memory Computing
  • Model Deployment
  • Network-on-Chip
  • Neural Network
  • Tensor Parallelism

Fingerprint

Dive into the research topics of 'BBP-DNS: Batch-Block Parallelism and Dual-NoC Scheduling for Accelerating GPT on Edge Devices'. Together they form a unique fingerprint.

Cite this