SKT:透過經驗證的合成數據生成進行大規模技能使用訓練
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
August 3, 2026
作者: Zelin Tan, Yiqun Zhang, Hao Li, Zhiyao Cui, Hejia Geng, Shao Zhang, Hangfan Zhang, Yang Chen, Xiaosong Wang, Lilong Wang, Zhenfei Yin, Shuyue Hu, Chen Zhang, Lei Bai
cs.AI
摘要
代理技能已成為賦予語言模型代理可重用程序性知識的重要機制。然而,僅提供技能並不能保證當前模型能有效地識別、應用和協調這些技能。為了提升技能使用能力,我們提出了 SKT——一個經過驗證的資料合成流程,能從大量代理技能中建構基於技能的任務與可執行的軌跡。SKT 會選擇合適的單技能與多技能配置,透過基於規則與基於代理的驗證並搭配回饋引導的修復來合成任務,且僅保留充分使用了每項所需技能的成功軌跡。利用 2,000 個公開技能,SKT 產生了 4,000 個任務包與 27,164 條經驗證的軌跡。基於相同的流程和一個互不重疊的測試池,我們進一步建構了 SkillEval——一個用於評估技能使用的保留式可執行基準。跨多種模型、基準和代理框架的實驗顯示,在 SKT 生成的軌跡上進行監督式微調能持續提升技能使用表現。驗證消融實驗、跨框架評估與規模化實驗進一步證明,這些提升依賴於高品質的監督,其效益超越單一代理介面,並會隨技能覆蓋範圍擴大而增加。綜合而言,這些結果確立了經驗證的資料合成作為一種有效且可擴展的技能使用訓練方法。
English
Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that current models can effectively identify, apply, and coordinate them. To improve skill-use capabilities, we introduce SKT, a verified data synthesis pipeline that constructs skill-grounded tasks and executable trajectories from large collections of agent skills. SKT selects suitable single-skill and multi-skill configurations, synthesizes tasks through rule-based and agent-based verification with feedback-guided repair, and retains only successful trajectories that substantially use every required skill. Using 2,000 public skills, SKT produces 4,000 task packages and 27,164 verified trajectories. Based on the same pipeline and a disjoint test pool, we further construct SkillEval, a held-out executable benchmark for evaluating skill use. Experiments across diverse models, benchmarks, and agent harnesses show that supervised fine-tuning on SKT-generated trajectories consistently improves skill-use performance. Verification ablations, cross-harness evaluation, and scaling experiments further demonstrate that these gains depend on high-quality supervision, extend beyond a single agent interface, and increase with broader skill coverage. Together, these results establish verified data synthesis as an effective and scalable approach for skill-use training.