DART-SD:マルチターン・ツール呼び出しエージェントの自己蒸留のためのダイヤモンドトポロジー対応検索とチューニング

DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents

August 19, 2026
著者: Hangrui Xu, Jiarui Wang, Yang Yang, Chuanbo Zhu, Fangda Chen, Ziqi Wu, Jingming Cai, Yan Song
cs.AI

要旨

大規模言語モデル(LLM)にマルチターンのツール呼び出し能力を備えることは、自律エージェントの構築に不可欠である。しかしながら、その進歩は、全長軌跡の模倣への依存によって根本的に制約されている。複数の順序非依存のサブゴールを伴うタスクでは、最適解空間は巨大な組み合わせ的ダイヤモンド格子を形成する。この豊かなトポロジーを一枚岩的な軌跡へと強制することは、深刻なトポロジカル崩壊を引き起こし、有効な代替探索を無差別に penalize して、方策の多様性を著しく低下させる。この問題に対処するため、我々は DART-SD(Diamond-topology Aware Retrieval and Tuning for Self-Distillation)を提案する。これは、グローバルな強制からトポロジー誘導型の局所的修正へとパラダイムを転換する新しいフレームワークである。DART-SD はまず、実行プロセスを収束する Interaction-State Transition Graph(ISTG)としてモデル化し、成功および失敗した探索経路が持つ本来のダイヤモンド構造を忠実に捕捉する。自律的ロールアウト中に、フレームワークは Critical Topological Breakpoint(CTB)を特定し、成功によって裏付けられた回復参照を取得する。最後に、CTB 誘導型の局所的監督を通じて漸進的な自己蒸留パラダイムを導入し、生成された回復ステップに対してのみ学習損失が計算されることを保証するとともに、有効な推論プレフィックスを破壊的な勾配更新から厳密に保護する。複雑なマルチターンツール呼び出しベンチマークにおける実験は、DART-SD が従来の全軌跡ベースラインを大幅に上回ることを実証している。
English
Equipping Large Language Models (LLMs) with multi-turn tool-calling capabilities is essential for building autonomous agents. However, progress is fundamentally limited by the reliance on full-length trajectory imitation. For tasks involving multiple order-independent sub-goals, the optimal solution space forms a vast combinatorial diamond lattice. Forcing this rich topology into monolithic trajectories causes a severe topological collapse, indiscriminately penalizing valid alternative explorations and severely degrading policy diversity. To address this, we propose DART-SD (Diamond-topology Aware Retrieval and Tuning for Self-Distillation), a novel framework that shifts the paradigm from global forcing to topology-guided localized correction. DART-SD first models the execution process as a converging Interaction-State Transition Graph (ISTG), faithfully capturing the inherent diamond topology of successful and failed exploratory paths. During autonomous rollouts, the framework identifies the Critical Topological Breakpoint (CTB) and retrieves success-supported recovery references. Finally, we introduce a progressive self-distillation paradigm through CTB-guided localized supervision, ensuring that the training loss is calculated exclusively on the generated recovery steps while strictly protecting the valid reasoning prefix from destructive gradient updates. Experiments on complex multi-turn tool-calling benchmarks demonstrate that DART-SD significantly outperforms traditional full-trajectory baselines.
PDF883September 1, 2026