Improving Inference Time of Deep Learning Model with Partial Skip of ReLU-fused Matrix Multiplication Operations

  • Kim, Sungkyun
  • Kim, Jaemin
  • Kim, Nahun
  • Kang, Mincheal
  • Seo, Jiwon
Citations

WEB OF SCIENCE

1
Citations

SCOPUS

3

초록

Deep learning has been expanding its application, while large-scale models tend to perform well. However, as such a model inevitably requires a vast amount of resources and computations, lengthy inference time is a crucial, but essential, consequence that needs to be optimized for the efficient utilization of deep learning. To achieve the goal, we aim at fusing the Rectified Linear Unit and matrix multiplication in the inference process, which we may reduce the total amount of computation by predicting the sign bit of output value. We propose four methods of prediction and statistically choose an optimal method for reducing inference time with low accuracy loss. © 2022 IEEE.

키워드

deep learning optimizationfully-connected layerinference optimizationomitted computationDeep learningDeep learning optimizationFully-connected layerInference optimizationITS applicationsLarge-scale modelingLearning modelsLearning optimizationsMatrix multiplication operationOmitted computationOptimisationsMatrix algebra
제목
Improving Inference Time of Deep Learning Model with Partial Skip of ReLU-fused Matrix Multiplication Operations
저자
Kim, SungkyunKim, JaeminKim, NahunKang, MinchealSeo, Jiwon
DOI
10.1109/ICEIC54506.2022.9748210
발행일
2022-04
유형
Proceedings Paper
저널명
2022 INTERNATIONAL CONFERENCE ON ELECTRONICS, INFORMATION, AND COMMUNICATION (ICEIC)
페이지
1 ~ 4