상세 보기
Hope: An Efficient Accelerator With Head-Wise Overlap Processing for Sparse Attention in Vision Transformer
- Heo, Jihyeon;
- Rho, Soomin;
- Kim, Kwangrae;
- Chung, Ki-Seok
WEB OF SCIENCE
0SCOPUS
1초록
Vision Transformers (ViTs) have revolutionized computer vision tasks with superior performance, but their quadratic computational complexity in self-attention remains a critical bottleneck. Sparse attention mitigates this issue by avoiding unnecessary computations, but existing accelerators enforce the sequential execution of QKV generation and sparse attention computations, which hinders execution time improvement. In this paper, we propose HOPE, a novel sparse attention accelerator for ViTs. HOPE accelerator introduces a head-wise overlap scheduling method and a new vector processing unit called SaVPU. The SaVPU augments conventional non-linear operation units with minimal hardware overhead to efficiently process sparse attention, such as SDDMM and SpMM operations. The head-wise overlap scheduling method enables concurrent execution of QKV generation and sparse attention computations. This overlapping execution significantly reduces latency. Extensive experiments on ViT models with sparsity levels between 60 % and 90 % demonstrate that HOPE achieves up to a 1.4× speedup (1.2× on average) over the state-of-the-art ViTCoD accelerator while incurring only 1.25 % additional hardware area. These results confirm the HOPE accelerator achieves significant speedup with minimal hardware overhead.
키워드
- 제목
- Hope: An Efficient Accelerator With Head-Wise Overlap Processing for Sparse Attention in Vision Transformer
- 저자
- Heo, Jihyeon; Rho, Soomin; Kim, Kwangrae; Chung, Ki-Seok
- 발행일
- 2025-06
- 유형
- Proceedings Paper
- 저널명
- 2025 IEEE SYMPOSIUM ON LOW-POWER AND HIGH-SPEED CHIPS AND SYSTEMS, COOL CHIPS
- 페이지
- 1 ~ 6