DSpark: 반자기회귀 생성을 활용한 신뢰도 스케줄링 기반 추측 디코딩
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
July 6, 2026
저자: Xin Cheng, Xingkai Yu, Chenze Shao, Jiashi Li, Yunfan Xiong, Yi Qian, Jiaqi Zhu, Shirong Ma, Xiaokang Zhang, Jiasheng Ye, Qinyu Chen, Chengqi Deng, Jiping Yu, Damai Dai, Zhengyan Zhang, Yixuan Wei, Yixuan Tan, Wenkai Yang, Runxin Xu, Yu Wu, Zhean Xu, Xuanyu Wang, Muyang Chen, Rui Tian, Xiao Bi, Zhewen Hao, Shaoyuan Chen, Huanqi Cao, Wentao Zhang, Anyi Xu, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang
cs.AI
초록
투기적 디코딩은 드래프트 생성을 타겟 검증에서 분리함으로써 대형 언어 모델(LLM) 추론을 가속화한다. 최근의 병렬 드래프터는 단일 순방향 패스에서 긴 토큰 시퀀스를 효율적으로 제안하지만, 토큰 간 의존성 부족으로 인해 급격한 수용률 감소를 겪는다. 더욱이, 이러한 확장된 블록을 무차별적으로 검증하는 것은 위험도가 높은 토큰에 중요한 배치 용량을 낭비하여, 높은 동시성을 가진 서빙 시스템에서 처리량을 심각하게 저하시킨다. 본 논문에서는 높은 처리량의 병렬 생성과 적응형 부하 인식 검증을 통합한 투기적 디코딩 프레임워크인 DSpark를 소개한다. 드래프트 품질을 유지하기 위해 DSpark는 반자기회귀 아키텍처를 활용하여, 병렬 백본을 경량 순차 모듈과 결합함으로써 블록 내 의존성 모델링을 도입하고 접미사 감쇠를 완화한다. 시스템 효율성을 최적화하기 위해 DSpark는 신뢰도 기반 검증을 사용하여, 각 요청에 대해 추정된 접두사 생존 확률과 엔진별 처리량 프로파일을 기반으로 검증 길이를 동적으로 조정한다. 다양한 도메인의 오프라인 벤치마크에서 DSpark는 최신 자기회귀 및 병렬 드래프터에 비해 수용된 길이를 크게 향상시킨다. 실시간 사용자 트래픽 하에서 DeepSeek-V4 서빙 시스템에 배포되었을 때, DSpark는 검증 낭비를 성공적으로 완화한다. 기존 프로덕션 기준인 MTP-1과 비교하여, DSpark는 동일한 처리량 수준에서 사용자당 생성 속도를 60~85% 가속화한다. 더욱 중요하게는, 엄격한 상호작용성 제약 조건에서 심각한 처리량 저하를 방지함으로써, 이전에는 달성할 수 없었던 성능 계층을 가능하게 하여 서빙 시스템의 파레토 경계를 이동시킨다.
English
Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification. While recent parallel drafters efficiently propose long token sequences in a single forward pass, they suffer from rapid acceptance decay due to a lack of inter-token dependencies. Furthermore, indiscriminately verifying these extended blocks wastes critical batch capacity on tokens with high rejection risks, severely degrading throughput in high-concurrency serving systems. We introduce DSpark, a speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification. To maintain draft quality, DSpark utilizes a semi-autoregressive architecture, coupling a parallel backbone with a lightweight sequential module, to introduce intra-block dependency modeling and mitigate suffix decay. To optimize system efficiency, DSpark employs confidence-scheduled verification, dynamically tailoring the verification length for each request based on estimated prefix survival probabilities and engine-specific throughput profiles. On offline benchmarks across diverse domains, DSpark substantially improves the accepted length over state-of-the-art autoregressive and parallel drafters. When deployed within the DeepSeek-V4 serving system under live user traffic, DSpark successfully mitigates verification waste. Compared to the established production baseline (MTP-1), DSpark accelerates per-user generation speeds by 60 to 85 percent at matched throughput levels. More importantly, by preventing severe throughput degradation under strict interactivity constraints, it enables performance tiers that were previously unattainable, shifting the Pareto frontier of our serving system.