NN-LUT: Neural Approximation of Non-Linear Operations for Efficient Transformer Inference

  • Yu, Joonsang
  • Park, Junki
  • 박성민
  • 김민수
  • Lee, Sihwa
  • ... Choi, Jungwook
  • 외 1명
Citations

WEB OF SCIENCE

70
Citations

SCOPUS

84

초록

Non-linear operations such as GELU, Layer normalization, and Soft-max are essential yet costly building blocks of Transformer models. Several prior works simplified these operations with look-up tables or integer computations, but such approximations suffer inferior accuracy or considerable hardware cost with long latency. This paper proposes an accurate and hardware-friendly approximation framework for efficient Transformer inference. Our framework employs a simple neural network as a universal approximator with its structure equivalently transformed into a Look-up table(LUT). The proposed framework called Neural network generated LUT(NN-LUT) can accurately replace all the non-linear operations in popular BERT models with significant reductions in area, power consumption, and latency.

키워드

look-up tableneural networknon-linear functiontransformerFunctionsTable lookupBuilding blockesHardware costLinear operationsLookup tables (LUTs)Neural-networksNon linearNonlinear functionsNormalisationTransformerTransformer modeling
제목
NN-LUT: Neural Approximation of Non-Linear Operations for Efficient Transformer Inference
저자
Yu, JoonsangPark, Junki박성민김민수Lee, SihwaLee, Dong HyunChoi, Jungwook
DOI
10.1145/3489517.3530505
발행일
2022-07
유형
Proceedings Paper
저널명
Proceedings - Design Automation Conference
페이지
577 ~ 582