Implementation of a CNN accelerator on an Embedded SoC Platform using SDSoC

초록

Today, Convolution Neural Networks (CNN) is adopted by various application areas such as computer vision, speech recognition, and natural language processing. Due to a massive amount of computing for CNN, CNN running on an embedded platform may not meet the performance requirement. In this paper, we propose a system-on-chip (SoC) CNN architecture synthesized by high level synthesis (HLS). HLS is an effective hardware (HW) synthesis method in terms of both development effort and performance. However, the implementation should be optimized carefully in order to achieve a satisfactory performance. Thus, we apply several optimization techniques to the proposed CNN architecture to satisfy the performance requirement. The proposed CNN architecture implemented on a Xilinx's Zynq platform has achieved 23% faster and 9.05 times better throughput per energy consumption than an implementation on an Intel i7 Core processor.

제목
Implementation of a CNN accelerator on an Embedded SoC Platform using SDSoC
저자
Sang-Soo, ParkKyeong-Bin, ParkChung, Ki Seok
DOI
10.1145/3193025.3193041
발행일
2018-02
유형
Proceeding
저널명
Proceedings of the 2nd International Conference on Digital Signal Processing
페이지
161 ~ 165