상세 보기
초록
Deep learning has been expanding its application, while large-scale models tend to perform well. However, as such a model inevitably requires a vast amount of resources and computations, lengthy inference time is a crucial, but essential, consequence that needs to be optimized for the efficient utilization of deep learning. To achieve the goal, we aim at fusing the Rectified Linear Unit and matrix multiplication in the inference process, which we may reduce the total amount of computation by predicting the sign bit of output value. We propose four methods of prediction and statistically choose an optimal method for reducing inference time with low accuracy loss. © 2022 IEEE.
키워드
deep learning optimization; fully-connected layer; inference optimization; omitted computation; Deep learning; Deep learning optimization; Fully-connected layer; Inference optimization; ITS applications; Large-scale modeling; Learning models; Learning optimizations; Matrix multiplication operation; Omitted computation; Optimisations; Matrix algebra
- 제목
- Improving Inference Time of Deep Learning Model with Partial Skip of ReLU-fused Matrix Multiplication Operations
- 저자
- Kim, Sungkyun; Kim, Jaemin; Kim, Nahun; Kang, Mincheal; Seo, Jiwon
- 발행일
- 2022-04
- 유형
- Proceedings Paper
- 저널명
- 2022 INTERNATIONAL CONFERENCE ON ELECTRONICS, INFORMATION, AND COMMUNICATION (ICEIC)
- 페이지
- 1 ~ 4