DSpark: 信頼度スケジュール投機的復号と半自己回帰生成
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
July 6, 2026
著者: Xin Cheng, Xingkai Yu, Chenze Shao, Jiashi Li, Yunfan Xiong, Yi Qian, Jiaqi Zhu, Shirong Ma, Xiaokang Zhang, Jiasheng Ye, Qinyu Chen, Chengqi Deng, Jiping Yu, Damai Dai, Zhengyan Zhang, Yixuan Wei, Yixuan Tan, Wenkai Yang, Runxin Xu, Yu Wu, Zhean Xu, Xuanyu Wang, Muyang Chen, Rui Tian, Xiao Bi, Zhewen Hao, Shaoyuan Chen, Huanqi Cao, Wentao Zhang, Anyi Xu, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang
cs.AI
要旨
推測デコーディングは、ドラフト生成をターゲット検証から分離することで、大規模言語モデル(LLM)の推論を高速化する。近年の並列ドラフターは、単一の順伝搬で長いトークン系列を効率的に提案できるが、トークン間の依存関係が欠如しているため、受入れ率が急速に低下するという問題を抱えている。さらに、これらの拡張ブロックを無差別に検証すると、拒否リスクの高いトークンにクリティカルなバッチ容量を浪費し、高並列サーバシステムにおけるスループットを著しく劣化させる。本稿では、高スループットな並列生成と適応的かつ負荷を考慮した検証を統合する推測デコーディングフレームワーク「DSpark」を提案する。DSparkは、ドラフト品質を維持するため、並列バックボーンと軽量な逐次モジュールを結合した半自己回帰アーキテクチャを採用し、ブロック内の依存関係モデリングを導入してサフィックス劣化を緩和する。システム効率を最適化するため、DSparkは信頼度スケジューリング検証を採用し、推定されたプレフィックス生存確率とエンジン固有のスループットプロファイルに基づいて、各リクエストの検証長を動的に調整する。多様なドメインにおけるオフラインベンチマークでは、DSparkは最先端の自己回帰型および並列型ドラフターを大幅に上回る受入れ長を達成する。実ユーザトラフィック下のDeepSeek-V4サーバシステムに導入した場合、DSparkは検証の無駄を効果的に削減する。確立されたプロダクションベースライン(MTP-1)と比較して、DSparkは同等スループットレベルでユーザあたりの生成速度を60~85%向上させる。さらに重要なことに、厳格な対話性制約下での深刻なスループット低下を防止することで、これまで達成不可能だったパフォーマンス層を実現し、サーバシステムのパレート最適フロンティアを押し広げる。
English
Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification. While recent parallel drafters efficiently propose long token sequences in a single forward pass, they suffer from rapid acceptance decay due to a lack of inter-token dependencies. Furthermore, indiscriminately verifying these extended blocks wastes critical batch capacity on tokens with high rejection risks, severely degrading throughput in high-concurrency serving systems. We introduce DSpark, a speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification. To maintain draft quality, DSpark utilizes a semi-autoregressive architecture, coupling a parallel backbone with a lightweight sequential module, to introduce intra-block dependency modeling and mitigate suffix decay. To optimize system efficiency, DSpark employs confidence-scheduled verification, dynamically tailoring the verification length for each request based on estimated prefix survival probabilities and engine-specific throughput profiles. On offline benchmarks across diverse domains, DSpark substantially improves the accepted length over state-of-the-art autoregressive and parallel drafters. When deployed within the DeepSeek-V4 serving system under live user traffic, DSpark successfully mitigates verification waste. Compared to the established production baseline (MTP-1), DSpark accelerates per-user generation speeds by 60 to 85 percent at matched throughput levels. More importantly, by preventing severe throughput degradation under strict interactivity constraints, it enables performance tiers that were previously unattainable, shifting the Pareto frontier of our serving system.