Hope: An Efficient Accelerator With Head-Wise Overlap Processing for Sparse Attention in Vision Transformer

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

1

초록

Vision Transformers (ViTs) have revolutionized computer vision tasks with superior performance, but their quadratic computational complexity in self-attention remains a critical bottleneck. Sparse attention mitigates this issue by avoiding unnecessary computations, but existing accelerators enforce the sequential execution of QKV generation and sparse attention computations, which hinders execution time improvement. In this paper, we propose HOPE, a novel sparse attention accelerator for ViTs. HOPE accelerator introduces a head-wise overlap scheduling method and a new vector processing unit called SaVPU. The SaVPU augments conventional non-linear operation units with minimal hardware overhead to efficiently process sparse attention, such as SDDMM and SpMM operations. The head-wise overlap scheduling method enables concurrent execution of QKV generation and sparse attention computations. This overlapping execution significantly reduces latency. Extensive experiments on ViT models with sparsity levels between 60 % and 90 % demonstrate that HOPE achieves up to a 1.4× speedup (1.2× on average) over the state-of-the-art ViTCoD accelerator while incurring only 1.25 % additional hardware area. These results confirm the HOPE accelerator achieves significant speedup with minimal hardware overhead.

키워드

Hardware AcceleratorSparse AttentionVision TransformerArray processingComputer hardwareComputer visionConcurrency controlParticle acceleratorsProgram processors
제목
Hope: An Efficient Accelerator With Head-Wise Overlap Processing for Sparse Attention in Vision Transformer
저자
Heo, JihyeonRho, SoominKim, KwangraeChung, Ki-Seok
DOI
10.1109/COOLCHIPS65488.2025.11018597
발행일
2025-06
유형
Proceedings Paper
저널명
2025 IEEE SYMPOSIUM ON LOW-POWER AND HIGH-SPEED CHIPS AND SYSTEMS, COOL CHIPS
페이지
1 ~ 6