상세 보기
Layer-wise Pruning of Transformer Attention Heads for Efficient Language Modeling
- Shim, Kyuhong;
- Choi, Iksoo;
- Sung, Wonyong;
- Choi, Jung wook
WEB OF SCIENCE
8SCOPUS
12초록
Recently, the necessity of multiple attention heads in transformer architecture has been questioned [1]. Removing less important heads from a large network is a promising strategy to reduce computation cost and parameters. However, pruning out attention heads in multihead attention does not evenly reduce the overall load, because feedforward modules are not affected. In this study, we apply attention head pruning on All-Attention [2] transformer, where savings in the computation are proportional to the number of pruned heads. This improved computing efficiency comes at the cost of pruning sensitivity, which we stabilize with three training techniques. Our attention head pruning enables a considerably fewer number of parameters with a comparable perplexity for transformer-based language modeling.
키워드
- 제목
- Layer-wise Pruning of Transformer Attention Heads for Efficient Language Modeling
- 저자
- Shim, Kyuhong; Choi, Iksoo; Sung, Wonyong; Choi, Jung wook
- 발행일
- 2021-11
- 유형
- Proceedings Paper
- 저널명
- 18TH INTERNATIONAL SOC DESIGN CONFERENCE 2021 (ISOCC 2021)
- 페이지
- 357 ~ 358