A Coalesced Tensor Reduction Architecture for Scalable All-Bank PIM Execution

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

The embedding layer in deep learning recommendation models (DLRM) is highly memory-bound and exhibits skewed, irregular access patterns. These characteristics lead to severe load imbalance and performance bottlenecks in processing in memory (PIM) architectures. We propose TRAM (Two-level Reduction Accelerator for Memory), a heterogeneous accelerator that integrates High Bandwidth Memory based PIM architecture (HBM-PIM) with conventional dual in-line memory modules (DIMMs) to accelerate batched embedding vector reductions. TRAM reduces redundant hot-vector accesses and employs a host-side scheduling mechanism that overlaps bank-PIM operations inside DRAM banks with logic-PIM operations, where processing units are located in the buffer die. This overlap eliminates command-bandwidth stalls and compute-bound delays. In addition, metadata-aware optimizations reduce row/column access overhead by reusing contiguous address patterns within each bank. Evaluation on six recommendation datasets and three embedding dimensions demonstrates that TRAM achieves up to 2.8× speedup and 3.0× energy reduction compared to state-of-the-art heterogeneous memory systems, while preserving full compatibility with the standard DRAM interface.

키워드

VectorsBandwidthComputer architectureRandom access memoryThroughputCircuits and systemsTensorsRecommender systemsMemory architectureFeature extractionRecommendation system3D-stacked memoryprocessing-in-memoryall-bank modebuffer dieArchitectureCost reductionDynamic random access storageEmbeddingsInterface statesMemory architectureThree dimensional integrated circuits
제목
A Coalesced Tensor Reduction Architecture for Scalable All-Bank PIM Execution
저자
Park, TaehyungLee, Hyuk-JaeRhee, Chae Eun
DOI
10.1109/JETCAS.2026.3657823
발행일
2026-06
유형
Article
저널명
IEEE Journal on Emerging and Selected Topics in Circuits and Systems
16
2
페이지
389 ~ 402