Implicit Jacobian Regularization Weighted with Impurity of Probability Output

Citations

SCOPUS

1

초록

The success of deep learning is greatly attributed to stochastic gradient descent (SGD), yet it remains unclear how SGD finds well-generalized models. We demonstrate that SGD has an implicit regularization effect on the logit-weight Jacobian norm of neural networks. This regularization effect is weighted with the impurity of the probability output, and thus it is active in a certain phase of training. Moreover, based on these findings, we propose a novel optimization method that explicitly regularizes the Jacobian norm, which leads to similar performance as other state-of-the-art sharpness-aware optimization methods.

키워드

Deep learningGradient methodsOptimizationStochastic models
제목
Implicit Jacobian Regularization Weighted with Impurity of Probability Output
저자
Lee, SungyoonPark, JinseongLee, Jaewook
발행일
2023-07
유형
Conference paper
저널명
Proceedings of Machine Learning Research (PMLR)
202
페이지
19094 ~ 19140