Finding Optimal Numerical Format for Sub-8-Bit Post-Training Quantization of Vision Transformers

Citations

SCOPUS

2

초록

Vision Transformers (ViTs) have gained significant attention for their exceptional model accuracies on computer vision applications, but their demanding memory requirements and computational complexity have hindered active deployment. Post-training quantization (PTQ) is a practical method to tackle this challenge by directly reducing ViT's bit-precision. However, diverse data characteristics across different operations of ViT cannot be well captured solely by a single numerical format (fixed or floating-point). This work proposes an analytical framework that optimizes the numerical format of each matrix multiplication of ViTs for mixed-format sub-8bit quantization. The extensive evaluation demonstrates that the proposed method can reduce the PTQ error and achieve state-of-the-art accuracy for popular ViT models.

키워드

fixed-pointfloating-pointPost-training quantizationvision Transformer
제목
Finding Optimal Numerical Format for Sub-8-Bit Post-Training Quantization of Vision Transformers
저자
이장환Hwang, YoungdeokChoi, Jungwook
DOI
10.1109/ICASSP49357.2023.10096798
발행일
2023-06
유형
Conference paper
저널명
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
페이지
1 ~ 5