DSpark:基於置信度排程的半自迴歸推測解碼
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
July 6, 2026
作者: Xin Cheng, Xingkai Yu, Chenze Shao, Jiashi Li, Yunfan Xiong, Yi Qian, Jiaqi Zhu, Shirong Ma, Xiaokang Zhang, Jiasheng Ye, Qinyu Chen, Chengqi Deng, Jiping Yu, Damai Dai, Zhengyan Zhang, Yixuan Wei, Yixuan Tan, Wenkai Yang, Runxin Xu, Yu Wu, Zhean Xu, Xuanyu Wang, Muyang Chen, Rui Tian, Xiao Bi, Zhewen Hao, Shaoyuan Chen, Huanqi Cao, Wentao Zhang, Anyi Xu, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang
cs.AI
摘要
推测解码通过将草稿生成与目标验证解耦,加速了大语言模型(LLM)的推理过程。尽管近期提出的并行草稿生成器能够在单次前向传播中高效生成较长的令牌序列,但由于缺乏令牌间的依赖关系,这些序列的接受率会快速衰减。此外,不加区分地验证这些扩展块,会浪费关键批量容量来处理具有高拒绝风险的令牌,从而严重降低高并发服务系统的吞吐量。我们提出了DSpark,一种将高吞吐量并行生成与自适应、负载感知验证相统一的推测解码框架。为了维持草稿质量,DSpark采用半自回归架构,将并行主干与轻量级顺序模块相结合,引入块内依赖建模,缓解后缀衰减问题。为了优化系统效率,DSpark采用置信度调度验证,根据估计的前缀存活概率以及引擎特有的吞吐量特征,动态调整每个请求的验证长度。在跨多种领域的离线基准测试中,与最先进的自回归和并行草稿生成器相比,DSpark显著提升了接受的令牌长度。当部署在DeepSeek-V4服务系统中并面对真实用户流量时,DSpark成功缓解了验证浪费问题。与已投产的基准(MTP-1)相比,DSpark在相同吞吐量水平下将每位用户的生成速度提升了60%到85%。更重要的是,通过防止在严格交互性约束下出现严重的吞吐量下降,DSpark实现了此前难以企及的性能档次,推动了我们服务系统的帕累托前沿发生位移。
English
Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification. While recent parallel drafters efficiently propose long token sequences in a single forward pass, they suffer from rapid acceptance decay due to a lack of inter-token dependencies. Furthermore, indiscriminately verifying these extended blocks wastes critical batch capacity on tokens with high rejection risks, severely degrading throughput in high-concurrency serving systems. We introduce DSpark, a speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification. To maintain draft quality, DSpark utilizes a semi-autoregressive architecture, coupling a parallel backbone with a lightweight sequential module, to introduce intra-block dependency modeling and mitigate suffix decay. To optimize system efficiency, DSpark employs confidence-scheduled verification, dynamically tailoring the verification length for each request based on estimated prefix survival probabilities and engine-specific throughput profiles. On offline benchmarks across diverse domains, DSpark substantially improves the accepted length over state-of-the-art autoregressive and parallel drafters. When deployed within the DeepSeek-V4 serving system under live user traffic, DSpark successfully mitigates verification waste. Compared to the established production baseline (MTP-1), DSpark accelerates per-user generation speeds by 60 to 85 percent at matched throughput levels. More importantly, by preventing severe throughput degradation under strict interactivity constraints, it enables performance tiers that were previously unattainable, shifting the Pareto frontier of our serving system.