상세 보기
A Coalesced Tensor Reduction Architecture for Scalable All-Bank PIM Execution
- Park, Taehyung;
- Lee, Hyuk-Jae;
- Rhee, Chae Eun
WEB OF SCIENCE
0SCOPUS
0초록
The embedding layer in deep learning recommendation models (DLRM) is highly memory-bound and exhibits skewed, irregular access patterns. These characteristics lead to severe load imbalance and performance bottlenecks in processing in memory (PIM) architectures. We propose TRAM (Two-level Reduction Accelerator for Memory), a heterogeneous accelerator that integrates High Bandwidth Memory based PIM architecture (HBM-PIM) with conventional dual in-line memory modules (DIMMs) to accelerate batched embedding vector reductions. TRAM reduces redundant hot-vector accesses and employs a host-side scheduling mechanism that overlaps bank-PIM operations inside DRAM banks with logic-PIM operations, where processing units are located in the buffer die. This overlap eliminates command-bandwidth stalls and compute-bound delays. In addition, metadata-aware optimizations reduce row/column access overhead by reusing contiguous address patterns within each bank. Evaluation on six recommendation datasets and three embedding dimensions demonstrates that TRAM achieves up to 2.8× speedup and 3.0× energy reduction compared to state-of-the-art heterogeneous memory systems, while preserving full compatibility with the standard DRAM interface.
키워드
- 제목
- A Coalesced Tensor Reduction Architecture for Scalable All-Bank PIM Execution
- 저자
- Park, Taehyung; Lee, Hyuk-Jae; Rhee, Chae Eun
- 발행일
- 2026-06
- 유형
- Article
- 권
- 16
- 호
- 2
- 페이지
- 389 ~ 402