Runtime-Controlled Evaluation of Speculative Decoding Methods

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Reported Speculative decoding (SD) speedups are difficult to compare because methods are commonly evaluated with different serving runtimes and configurations. We present SpecLLM, a pluggable evaluation framework that applies common scheduling, batching, KV-cache, CUDA Graph, and execution-backend policies across methods while preserving method-required proposal and verification paths. We implement eight Speculative Decoding (SD) methods and evaluate them across batch sizes, controlled online loads, and two model scales. We decompose throughput speedup as =/, where is the average number of committed tokens per SD step and c is the dimensionless execution cost of that step relative to a matched autoregressive decode step. The measured c is conditional on the method–runtime interaction, hardware, workload, baseline, and configuration. Under the evaluated conditions, method rankings change with batch size and online load, and a longer commit length does not necessarily yield lower latency when execution cost or queueing increases. SpecLLM therefore provides controlled within-runtime comparability rather than runtime-independent fairness; SD results should report (,), latency, feasibility, and the conditions under which they were measured.

키워드

batchingbenchmarkingGPU runtime systemsinference accelerationKV cacheLLM servingspeculative decodingAccelerationDecodingProgram processorsScheduling algorithms
제목
Runtime-Controlled Evaluation of Speculative Decoding Methods
저자
Kim, SungkyunKim, JaeminCho, YeongpilSeo, Jiwon
DOI
10.3390/electronics15143179
발행일
2026-07
유형
Article
저널명
Electronics (Switzerland)
15
14
페이지
1 ~ 29

파일 다운로드