ARNorm: Hardware-Efficient Normalization for Lightweight Edge Models

Citations

SCOPUS

0

초록

We propose ARNorm, a hardware-friendly normalization algorithm designed for efficient inference in transformer-based models. ARNorm combines the structural simplicity of RMSNorm [1] with the hardware-optimized techniques of AILayerNorm, introduced in SOLE [2] achieving accurate normalization using only 8-bit integer precision (INT8) arithmetic. By employing dynamic compression and a priority encoder-based Look-Up Table (LUT) for root approximation, ARNorm eliminates costly floating-point operations such as mean, variance, and square root calculations. Experiments on six pre-trained Vision Transformer models demonstrate that ARNorm reduces quantization error by up to 10% compared to AILayerNorm and maintains accuracy comparable to 32-bit floating point precision(FP32)-based RMSNorm, making it highly suitable for edge and embedded AI applications.

키워드

AILayerNormHardware AcceleratorINT8Layer NormalizationQuantizationRMSNormTransformerVision Trans-formerComputation theoryComputer hardwareComputer visionImage codingInference enginesNumerical analysisTable lookup
제목
ARNorm: Hardware-Efficient Normalization for Lightweight Edge Models
저자
Kim, SunyeopRhee, Chae-eun
DOI
10.1109/ITC-CSCC66376.2025.11137703
발행일
2025-09
유형
Conference paper
저널명
2025 International Technical Conference on Circuits/Systems, Computers, and Communications, ITC-CSCC 2025
페이지
1 ~ 3