Optimizing Exponent Bias for Sub-8bit Floating-Point Inference of Fine-tuned Transformers

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

3

초록

The Transformer-based fine-tuned neural networks have demonstrated remarkable success in natural language processing (NLP) at the cost of a substantial computational burden. Post-training quantization (PTQ) is a promising technique to reduce the computational cost without expensive re-training. But prior works either demand complex calibration or suffer noticeable accuracy degradation. This paper proposes a practical method for sub-8bit floating-point (FP) PTQ. The proposed method optimizes the exponent bias to minimize quantization error in terms of signal-to-quantization noise ratio (SQNR) progressively like stochastic gradient descent. We evaluate that the proposed method achieves close to full-precision model accuracy for 6 to 8 bit FP PTQ of fine-tuned BERT on GLUE and SQuAD tasks with negligible run-time overhead.

키워드

BERTexponent biasfloating-pointpost-training quantizationreduced-precisionSQNRTransformerDigital arithmeticNatural language processing systemsOptimizationQuantization (signal)Stochastic systemsGradient methodsBERTExponent biasFloating pointsNatural languagesNeural-networksPost-training quantizationQuantisationReduced precisionSignal to quantization noise ratiosTransformer
제목
Optimizing Exponent Bias for Sub-8bit Floating-Point Inference of Fine-tuned Transformers
저자
이장환Choi, Jung wook
DOI
10.1109/AICAS54282.2022.9869965
발행일
2022-06
유형
Proceedings Paper
저널명
2022 IEEE INTERNATIONAL CONFERENCE ON ARTIFICIAL INTELLIGENCE CIRCUITS AND SYSTEMS (AICAS 2022): INTELLIGENT TECHNOLOGY IN THE POST-PANDEMIC ERA
페이지
98 ~ 101