A Low Power Attention and Softmax Accelerator for Large Language Models Inference

Citations

SCOPUS

4

초록

Transformer-based models, essential for high-performing Large Language Models (LLMs), surpass traditional Deep Neural Networks but require substantial computational resources. Therefore, more efficient transformer algorithms and accelerators are required to reduce the computational cost and power consumption of LLMs. We observed that as the sequence length increases, softmax operations, which are the key operation of the transformer self-attention mechanism, become the major bottleneck. In this paper, we propose Cross-Road Softmax, an optimized algorithm designed for the softmax operation within the attention layer, specifically tailored for inference in LLMs. Our software experiment was conducted on 8 Natural Language Processing benchmarks for evaluation. Furthermore, we design a Cross-Road Accel using the proposed Cross-Road Softmax that accelerates softmax function of the self-attention layer. We implement Cross-Road Accel in RTL and synthesize it with Syn-opsys Design Compiler using Nangate 15nm open cell library to obtain power and area statistics. In summary, on average, Cross-Road Accel achieves an approximately 3.5 × increase in energy efficiency compared to state-of-the-art transformer accelerators.

키워드

AI acceleratorAlgorithm-Hardware Co-DesignLLMsLow Power DesignNLPSoftmaxTransformerBenchmarkingIntegrated circuit designMultilayer neural networksPrinted circuit designProblem oriented languagesProgram compilersStructural dynamics
제목
A Low Power Attention and Softmax Accelerator for Large Language Models Inference
저자
Kim, Jeong-HyunKim, Chan-HoonRho, Soo-MinChung, Ki-Seok
DOI
10.1109/ICCE-Asia63397.2024.10773935
발행일
2024-12
유형
Conference paper
저널명
2024 IEEE International Conference on Consumer Electronics-Asia, ICCE-Asia 2024
페이지
1 ~ 4