TY - GEN
T1 - MoEA
T2 - 31st Asia and South Pacific Design Automation Conference, ASP-DAC 2026
AU - Dang, Qiwei
AU - Ma, Chengyu
AU - Huo, Zhiwang
AU - Yang, Guoming
AU - Xia, Tian
AU - Zhao, Wenzhe
AU - Ren, Pengju
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Recent vision models integrating Convolutional Neural Networks (CNNs) and attention-based Transformers have achieved unprecedented accuracy, but face significant obstacles for edge deployment: sensitivity to quantization formats (especially for nonlinear functions in Transformers), high computational instructions complexity, and the inability to perform on-device fine-tuning for distribution shifts. To address these challenges, we propose Mixture-of-Edge-Architectures(MoEA), a RISC-V-based accelerator with three key innovations. First, the Mixed Fixed/Floating-Point (MFP) format can unify INT8, INT16, Shared Exponent Floating-Point (SFP8), FP8, and FP16 into a single data path, supporting the optimal mixed-precision strategy. Second, the typical VLIW instruction is condensed into a single direct memory access instruction to improve the computation performance. Third, a specialized engine integrates an FPGA-optimized Matrix Processing Unit(MPU), a Single Instruction Multi Data(SIMD)-based Vector Processing Unit(VPU), and an enhanced Direct Memory Access(DMA) for back-propagation. Implemented on FPGAs, MoEA delivers 420 GOPS (0.99 GOPS/DSP) on ZCU102, with ResNet18 inference at 16 ms and fine-tuning at 70ms. An 8-cluster variant on XCVU9P reduces ViT-Base latency to 23ms, 1.81 × ∼ 2.13 × faster than prior accelerators, supporting versatile model deployment and adaptive edge intelligence.
AB - Recent vision models integrating Convolutional Neural Networks (CNNs) and attention-based Transformers have achieved unprecedented accuracy, but face significant obstacles for edge deployment: sensitivity to quantization formats (especially for nonlinear functions in Transformers), high computational instructions complexity, and the inability to perform on-device fine-tuning for distribution shifts. To address these challenges, we propose Mixture-of-Edge-Architectures(MoEA), a RISC-V-based accelerator with three key innovations. First, the Mixed Fixed/Floating-Point (MFP) format can unify INT8, INT16, Shared Exponent Floating-Point (SFP8), FP8, and FP16 into a single data path, supporting the optimal mixed-precision strategy. Second, the typical VLIW instruction is condensed into a single direct memory access instruction to improve the computation performance. Third, a specialized engine integrates an FPGA-optimized Matrix Processing Unit(MPU), a Single Instruction Multi Data(SIMD)-based Vector Processing Unit(VPU), and an enhanced Direct Memory Access(DMA) for back-propagation. Implemented on FPGAs, MoEA delivers 420 GOPS (0.99 GOPS/DSP) on ZCU102, with ResNet18 inference at 16 ms and fine-tuning at 70ms. An 8-cluster variant on XCVU9P reduces ViT-Base latency to 23ms, 1.81 × ∼ 2.13 × faster than prior accelerators, supporting versatile model deployment and adaptive edge intelligence.
KW - Deep Neural Network
KW - Hardware Accelerator
KW - Low-precision data format
KW - RISC-V
UR - https://www.scopus.com/pages/publications/105041715977
U2 - 10.1109/ASP-DAC66049.2026.11420594
DO - 10.1109/ASP-DAC66049.2026.11420594
M3 - 会议稿件
AN - SCOPUS:105041715977
T3 - Proceedings of the Asia and South Pacific Design Automation Conference, ASP-DAC
SP - 325
EP - 332
BT - ASP-DAC 2026 - 31st Asia and South Pacific Design Automation Conference, Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 19 January 2026 through 22 January 2026
ER -