상세 보기
Implicit Jacobian Regularization Weighted with Impurity of Probability Output
- Lee, Sungyoon;
- Park, Jinseong;
- Lee, Jaewook
Citations
SCOPUS
1초록
The success of deep learning is greatly attributed to stochastic gradient descent (SGD), yet it remains unclear how SGD finds well-generalized models. We demonstrate that SGD has an implicit regularization effect on the logit-weight Jacobian norm of neural networks. This regularization effect is weighted with the impurity of the probability output, and thus it is active in a certain phase of training. Moreover, based on these findings, we propose a novel optimization method that explicitly regularizes the Jacobian norm, which leads to similar performance as other state-of-the-art sharpness-aware optimization methods.
키워드
Deep learning; Gradient methods; Optimization; Stochastic models
- 제목
- Implicit Jacobian Regularization Weighted with Impurity of Probability Output
- 저자
- Lee, Sungyoon; Park, Jinseong; Lee, Jaewook
- 발행일
- 2023-07
- 유형
- Conference paper
- 저널명
- Proceedings of Machine Learning Research (PMLR)
- 권
- 202
- 페이지
- 19094 ~ 19140