Log-quantization on GRU networks

Citations

SCOPUS

0

초록

Today, recurrent neural network (RNN) is used in various applications like image captioning, speech recognition and machine translation. However, because of data dependencies, recurrent neural network is hard to parallelize. Furthermore, to increase network’s accuracy, recurrent neural network uses complicated cell units such as long short-term memory (LSTM) and gated recurrent unit (GRU). To run such models on an embedded system, the size of the network model and the amount of computation need to be reduced to achieve low power consumption and low required memory bandwidth. In this paper, implementation of RNN based on GRU with a logarithmic quantization method is proposed. The proposed implementation is synthesized using high-level synthesis (HLS) targeting Xilinx ZCU102 FPGA running at 100MHz. The proposed implementation with an 8-bit log-quantization achieves 90.57% accuracy without re-training or fine-tuning. And the memory usage is 31% lower than that for an implementation with 32-bit floating point data representation.

키워드

AICNNFPGAHLSHW/SW Co-DesignLeNet-5SDSoCArtificial intelligenceDigital arithmeticField programmable gate arrays (FPGA)Hardware-software codesignHigh level synthesisLogic SynthesisLow power electronicsSpeech recognitionSpeech transmissionData dependenciesFloating-point dataHW/SW CodesignLeNet-5Low-power consumptionMachine translationsRecurrent neural network (RNN)SDSoCLong short-term memory
제목
Log-quantization on GRU networks
저자
Park, Sang-Ki Park, Sang-SooChung, Ki Seok
DOI
10.1145/3290420.3290443
발행일
2018-11
유형
Conference Paper
저널명
ACM International Conference Proceeding Series
페이지
112 ~ 116