Power-Efficient Deep Neural Network Accelerator Minimizing Global Buffer Access without Data Transfer between Neighboring Multiplier-Accumulator Units

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

This paper presents a novel method for minimizing the power consumption of weight data movements required by a convolutional operation performed on a two-dimensional multiplier-accumulator (MAC) array of a deep neural-network accelerator. The proposed technique employs a local register file (LRF) at each MAC unit in a manner such that once weight pixels are read from the global buffer into the LRF, they are reused from the LRF as many times as desired instead of being repeatedly fetched from the global buffer in each convolutional operation. One of the most evident merits of the proposed method is that the procedure is completely free from the burden of data transfer between neighboring MAC units. It was found from our simulations that the proposed method provides a power saving of approximately 83.33% and 97.62% compared with the power savings recorded by the conventional methods, respectively, when the dimensions of the input data matrix and weight matrix are 128 x 128 and 5 x 5, respectively. The power savings increase as the dimensions of the input data matrix or weight matrix increase.

키워드

deep learning acceleratorfield-programmable gate array (FPGA)deep neural networks (DNNs)COPROCESSOR
제목
Power-Efficient Deep Neural Network Accelerator Minimizing Global Buffer Access without Data Transfer between Neighboring Multiplier-Accumulator Units
저자
Lee, JeonghyeokHan, SangwookChoi, SeungwonChoi, Jungwook
DOI
10.3390/electronics11131996
발행일
2022-07
유형
Article
저널명
ELECTRONICS
11
13
페이지
1 ~ 12

파일 다운로드