A 7-nm Four-Core Mixed-Precision AI Chip With 26.2-TFLOPS Hybrid-FP8 Training, 104.9-TOPS INT4 Inference, and Workload-Aware Throttling

  • Lee, Sae Kyu
  • Agrawal, Ankur
  • Silberman, Joel
  • Ziegler, Matthew
  • Kang, Mingu
  • ... Choi, Jungwook
  • 외 38명
Citations

WEB OF SCIENCE

32
Citations

SCOPUS

36

초록

Reduced precision computation is a key enabling factor for energy-efficient acceleration of deep learning (DL) applications. This article presents a 7-nm four-core mixed-precision artificial intelligence (AI) chip that supports four compute precisions--FP16, Hybrid-FP8 (HFP8), INT4, and INT2--to support diverse application demands for training and inference. The chip leverages cutting-edge algorithmic advances to demonstrate leading-edge power efficiency for 8-bit floating-point (FP8) training and INT4 inference without model accuracy degradation. A new HFP8 format combined with separation of the floating- and fixed-point pipelines and aggressive circuit/architecture optimization enables performance improvements while maintaining high compute utilization. A high-bandwidth ring protocol enables efficient data communication, while power management using workload-aware clock throttling maximizes performance within a given power budget. The AI chip demonstrates 3.58-TFLOPS/W peak energy efficiency and 26.2-TFLOPS peak performance for HFP8 iso-accuracy training, and 16.9-TOPS/W peak energy efficiency and 104.9-TOPS peak performance for INT4 iso-accuracy inference.

키워드

TrainingArtificial intelligenceAI acceleratorsInference algorithmsComputer architectureBandwidthSystem-on-chipApproximate computingartificial intelligence (AI)deep neural networks (DNNs)hardware acceleratorsmachine learning (ML)reduced precision computationPROCESSOR
제목
A 7-nm Four-Core Mixed-Precision AI Chip With 26.2-TFLOPS Hybrid-FP8 Training, 104.9-TOPS INT4 Inference, and Workload-Aware Throttling
저자
Lee, Sae KyuAgrawal, AnkurSilberman, JoelZiegler, MatthewKang, MinguVenkataramani, SwagathCao, NianzhengFleischer, BruceGuillorn, MichaelCohen, MatthewMueller, Silvia M.Oh, JinwookLutz, MartinJung, JinwookKoswatta, SiyuZhou, ChingZalani, VidhiKar, MonodeepBonanno, JamesCasatuta, RobertChen, Chia-YuChoi, JungwookHaynie, HowardHerbert, AlyssaJain, RadhikaKim, Kyu-HyounLi, YulongRen, ZhibinRider, ScotSchaal, MarcelSchelm, KerstinScheuermann, Michael R.Sun, XiaoTran, HungWang, NaigangWang, WeiZhang, XinShah, VinayCurran, BrianSrinivasan, VijayalakshmiLu, Pong-FeiShukla, SunilGopalakrishnan, KailashChang, Leland
DOI
10.1109/JSSC.2021.3120113
발행일
2022-01
유형
Article; Early Access
저널명
IEEE Journal of Solid-State Circuits
57
1
페이지
182 ~ 197