Layer-wise Pruning of Transformer Attention Heads for Efficient Language Modeling

Citations

WEB OF SCIENCE

8
Citations

SCOPUS

12

초록

Recently, the necessity of multiple attention heads in transformer architecture has been questioned [1]. Removing less important heads from a large network is a promising strategy to reduce computation cost and parameters. However, pruning out attention heads in multihead attention does not evenly reduce the overall load, because feedforward modules are not affected. In this study, we apply attention head pruning on All-Attention [2] transformer, where savings in the computation are proportional to the number of pruned heads. This improved computing efficiency comes at the cost of pruning sensitivity, which we stabilize with three training techniques. Our attention head pruning enables a considerably fewer number of parameters with a comparable perplexity for transformer-based language modeling.

키워드

multihead attentionpruningtransformerComputational linguisticsComputation costsComputing efficiencyFeed forwardLanguage modelLarger networksLayer-wiseMultiheadMultihead attentionPruningTransformerModeling languages
제목
Layer-wise Pruning of Transformer Attention Heads for Efficient Language Modeling
저자
Shim, KyuhongChoi, IksooSung, WonyongChoi, Jung wook
DOI
10.1109/ISOCC53507.2021.9613933
발행일
2021-11
유형
Proceedings Paper
저널명
18TH INTERNATIONAL SOC DESIGN CONFERENCE 2021 (ISOCC 2021)
페이지
357 ~ 358