상세 보기
Runtime-Controlled Evaluation of Speculative Decoding Methods
- Kim, Sungkyun;
- Kim, Jaemin;
- Cho, Yeongpil;
- Seo, Jiwon
WEB OF SCIENCE
0SCOPUS
0초록
Reported Speculative decoding (SD) speedups are difficult to compare because methods are commonly evaluated with different serving runtimes and configurations. We present SpecLLM, a pluggable evaluation framework that applies common scheduling, batching, KV-cache, CUDA Graph, and execution-backend policies across methods while preserving method-required proposal and verification paths. We implement eight Speculative Decoding (SD) methods and evaluate them across batch sizes, controlled online loads, and two model scales. We decompose throughput speedup as =/, where is the average number of committed tokens per SD step and c is the dimensionless execution cost of that step relative to a matched autoregressive decode step. The measured c is conditional on the method–runtime interaction, hardware, workload, baseline, and configuration. Under the evaluated conditions, method rankings change with batch size and online load, and a longer commit length does not necessarily yield lower latency when execution cost or queueing increases. SpecLLM therefore provides controlled within-runtime comparability rather than runtime-independent fairness; SD results should report (,), latency, feasibility, and the conditions under which they were measured.
키워드
- 제목
- Runtime-Controlled Evaluation of Speculative Decoding Methods
- 저자
- Kim, Sungkyun; Kim, Jaemin; Cho, Yeongpil; Seo, Jiwon
- 발행일
- 2026-07
- 유형
- Article
- 저널명
- Electronics (Switzerland)
- 권
- 15
- 호
- 14
- 페이지
- 1 ~ 29