Understanding and Reducing Weight-Load Overhead of Systolic Deep Learning Accelerators

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

As an energy-efficient computing engine for deep neural network inference, 2D systolic array architectures have been widely adopted in modern deep learning accelerators. However, despite high compute density and energy-efficient data passing, systolic accelerators suffer a non-Trivial overhead of loading data stationed inside their local register file. This loading overhead becomes a critical issue when a frequent reload of stationary data (e.g., weight parameters) is required. This paper proposes a simple yet practical SW-HW co-optimization that reverses the weight-load order and adds a dedicated path for weight-load. On diverse deep learning applications, the proposed method reduces the weight-load overhead and achieves up to 1.8× speedup with 40% energy savings.

키워드

acceleratordeep neural networksystolic arrayAccelerationComputer aided designCost reductionEnergy efficiencyLow power electronicsNetwork architectureSystolic arraysComputing enginesCritical issuesEnergy efficientEnergy-efficient computingLoading dataNetwork inferenceNon-trivialRegister filesSystolic array architectureWeight parametersDeep neural networks
제목
Understanding and Reducing Weight-Load Overhead of Systolic Deep Learning Accelerators
저자
Joo, JinWonYoon, MinyongChoi, Jung wookKang, MinguLee, JongGeonSo, JinInYun, IlKwonKwon, YongsukKim, KyungSoo
DOI
10.1109/ISOCC53507.2021.9613929
발행일
2021-11
유형
Proceedings Paper
저널명
18TH INTERNATIONAL SOC DESIGN CONFERENCE 2021 (ISOCC 2021)
페이지
413 ~ 414