A Scalable Dual-Input/Output Heap-Based Top-K Unit for Vector Search Accelerators

Citations

SCOPUS

0

초록

Modern vector search accelerators require efficient top-K maintenance to process massive similarity scores generated by distance computation units. As the number of candidates (K) increases, conventional architectures such as a systolic priority queue and a pipelined heap face a trade-off between throughput and hardware complexity, often becoming a primary bottleneck in large-scale search systems. To address this limitation, we propose a dual-input/output heap-based top-K unit that achieves high throughput with low hardware overhead. The proposed top-K unit processes two input elements simultaneously using a hierarchical sub-heap structure, enabling continuous dual-element insertion and extraction every two cycles. This design achieves logarithmic scaling in the number of required comparators with respect to K, enabling efficient operation for large-K configurations. Furthermore, a large portion of flip-flop-based storage is replaced with high-density SRAM at deeper heap levels, mitigating register overhead while preserving the top-K candidate set. Synthesized in a Samsung 28 nm CMOS process, the proposed architecture demonstrates strong scalability and hardware efficiency for large-K vector search accelerators.

키워드

Hardware AcceleratorHeap ArchitectureTop-K UnitVector SearchAccelerationFlip flop circuitsParticle acceleratorsSearch enginesStatic random access storageVectors
제목
A Scalable Dual-Input/Output Heap-Based Top-K Unit for Vector Search Accelerators
저자
Kim, Ji SooYang, HannahKim, SujinKim, Ji-Hoon
DOI
10.1109/ISCAS66217.2026.11562482
발행일
2026-06
유형
Conference Paper
저널명
Proceedings - IEEE International Symposium on Circuits and Systems
페이지
4038 ~ 4042