Convergence-aware neural network training

  • Oh, Hyungjun
  • Yu, Yongseung
  • Ryu, Giha
  • Ahn, Gunjoo
  • Jeong, Yuri
  • 외 2명
Citations

SCOPUS

4

초록

Training a deep neural network(DNN) is expensive, requiring a large amount of computation time. While the training overhead is high, not all computation in DNN training is equal. Some parameters converge faster and thus their gradient computation may contribute little to the parameter update; in nearstationary points a subset of parameters may change very little. In this paper we exploit the parameter convergence to optimize gradient computation in DNN training. We design a light-weight monitoring technique to track the parameter convergence; we prune the gradient computation stochastically for a group of semantically related parameters, exploiting their convergence correlations. These techniques are efficiently implemented in existing GPU kernels. In our evaluation the optimization techniques substantially and robustly improve the training throughput for four DNN models on three public datasets.

키워드

Computer aided designDeep neural networksParameter estimationComputation timeGradient computationMonitoring techniquesNeural network trainingOptimization techniquesParameter convergenceTraining overheadTraining throughputsNeural networks
제목
Convergence-aware neural network training
저자
Oh, Hyungjun Yu, YongseungRyu, GihaAhn, GunjooJeong, Yuri Park, YongjunSeo, Jiwon
DOI
10.1109/DAC18072.2020.9218518
발행일
2020-07
유형
Conference Paper
저널명
Proceedings - Design Automation Conference
2020-July
페이지
1 ~ 6