Lightweight Error Correction for In-Storage Acceleration of Large Language Model Inference

Citations

SCOPUS

2

초록

As large language models (LLMs) expand their sizes, conventional GPU-based LLM inference systems face memory bandwidth and capacity limitations. An LLM inference accelerator using NAND flash storage has been proposed to overcome these challenges. However, this necessitates a significant expansion of flash channels to ensure adequate bandwidth for inference, subsequently escalating error correction code (ECC) costs. This paper examines the impact of flash memory errors on LLM inference accuracy and explores the possibility of lightweight ECC by leveraging LLM's inherent error resilience. We analyze the impact of 1) high-order bit indices masking for FP32 LLM parameters, 2) clipping, and 3) a dependency by parameter type of error robustness, and show that a combination of them can reduce ECC bandwidth by up to 9.38%.

키워드

error correction codelarge language modelNAND flash errorsError correction codesErrors correctionInference systemsLanguage modelLarge language modelMemory bandwidthsMemory capacityModel inferenceNAND FlashNAND flash error
제목
Lightweight Error Correction for In-Storage Acceleration of Large Language Model Inference
저자
Jeong, JinwooAhn, ByungminShin, DongminChoi, Jungwook
DOI
10.1109/ICEIC61013.2024.10457117
발행일
2024-01
유형
Conference paper
저널명
2024 International Conference on Electronics, Information, and Communication, ICEIC 2024
페이지
1 ~ 4