An Area-Efficient Mixed-Precision Accelerator with Output-Error-Based Quantization for ViT

  • Park, Subin
  • Ahn, Juhyuk
  • Rho, Soomin
  • Kim, Kwangrae
  • Chung, Ki-Seok
Citations

SCOPUS

0

초록

Vision Transformer (ViT) has achieved remarkable performance in computer vision tasks. However, its large number of parameters poses challenges for deployment on resourceconstrained devices. Mixed-precision quantization is widely used to reduce the model size. To improve accuracy while minimizing the use of high bit-width precision, selecting the appropriate precision for each tensor is crucial. In this paper, we propose a precision selection strategy that leverages the mean squared error of linear operation outputs to improve accuracy with minimal use of high bit-width tensors. Moreover, we propose a processing element that shares most of its internal resources to support mixed precision. On ViT-Base with ImageNet, our method achieves a 0.706% accuracy improvement and 1.83× speedup over a prior work with identical area constraints.

키워드

Hardware AcceleratorMixed-Precision QuantizationVision TransformerComputer visionImage enhancementParticle acceleratorsTensors
제목
An Area-Efficient Mixed-Precision Accelerator with Output-Error-Based Quantization for ViT
저자
Park, SubinAhn, JuhyukRho, SoominKim, KwangraeChung, Ki-Seok
DOI
10.1109/ISOCC66390.2025.11330033
발행일
2026-01
유형
Conference paper
저널명
International SoC Design Conference 2025, ISOCC 2025 - Proceedings of Technical Papers
페이지
1 ~ 2