Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices

Citations

SCOPUS

0

초록

In Transformer models, non-GEMM (non-General Matrix Multiplication) operations—especially Softmax and Layer Normalization (LayerNorm)—often dominate hardware cost due to their nonlinear nature. To address this, previous approximation studies mainly target rank-oriented tasks, which is acceptable for classification. However, edge Natural Language Processing (NLP) applications and edge generative AI are largely evaluated based on score-oriented tasks, so normalization-guaranteed nonGEMM operations are essential. We propose a hardware-efficient Softmax and LayerNorm with Guaranteed Normalization for Edge devices. Our design employs hardware-efficient approximation methods while preserving the normalization (Softmax: Σp=1, LayerNorm: σ =1). Our architecture is described in Verilog HDL and synthesized using the Samsung 28nm CMOS process. In accuracy evaluation, we achieve high accuracy with minimal degradation; GLUE +0.07%, SQuAD -0.01%, perplexity -0.09%. Implementation results show that our architecture is small; 942µm2 for Softmax, 1199µm2 for LayerNorm. Compared to the state of the art, we achieve up to 11x and 14x reduction in area, respectively.

키워드

Edge DevicesHardwareLayer NormalizationNLPScore-OrientedSoftmaxTransformerApproximation theoryComputation theoryComputer hardware description languagesKnowledge based systemsNatural language processing systemsSignal processing
제목
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
저자
Choi, DawonKim, HanaKim, Ji-Hoon
DOI
10.1109/ISCAS66217.2026.11562304
발행일
2026-06
유형
Conference Paper
저널명
Proceedings - IEEE International Symposium on Circuits and Systems
페이지
1246 ~ 1250