SKT: 検証済み合成データ生成による大規模スキル活用トレーニング
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
August 3, 2026
著者: Zelin Tan, Yiqun Zhang, Hao Li, Zhiyao Cui, Hejia Geng, Shao Zhang, Hangfan Zhang, Yang Chen, Xiaosong Wang, Lilong Wang, Zhenfei Yin, Shuyue Hu, Chen Zhang, Lei Bai
cs.AI
要旨
エージェントスキルは、言語モデルエージェントに再利用可能な手続き的知識を備えるための重要なメカニズムとなっています。しかしながら、スキルを提供するだけでは、現在のモデルがそれらを効果的に特定し、適用し、連携させることができることは保証されません。このスキル使用能力を向上させるために、我々はSKTを導入します。SKTは、多数のエージェントスキルからスキルに基づくタスクと実行可能な軌跡を構築する、検証済みデータ合成パイプラインです。SKTは、適切な単一スキルおよび複数スキルの構成を選択し、ルールベースおよびエージェントベースの検証とフィードバック誘導型の修復を通じてタスクを合成し、必要なすべてのスキルを実質的に使用した成功軌跡のみを保持します。2,000の公開スキルを用いて、SKTは4,000のタスクパッケージと27,164の検証済み軌跡を生成します。同じパイプラインと重複しないテストプールに基づき、スキル使用を評価するためのホールドアウト実行可能ベンチマークであるSkillEvalも構築します。多様なモデル、ベンチマーク、エージェントハーネスにわたる実験により、SKTが生成した軌跡に対する教師ありファインチューニングが、スキル使用パフォーマンスを一貫して向上させることが示されました。検証アブレーション、クロスハーネス評価、スケーリング実験はさらに、これらの向上が高品質な教師信号に依存し、単一のエージェントインターフェースを超えて及ぶものであり、より広範なスキルカバレッジとともに増大することを実証しています。これらの結果は総合的に、検証済みデータ合成がスキル使用トレーニングのための効果的かつスケーラブルなアプローチであることを確立しています。
English
Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that current models can effectively identify, apply, and coordinate them. To improve skill-use capabilities, we introduce SKT, a verified data synthesis pipeline that constructs skill-grounded tasks and executable trajectories from large collections of agent skills. SKT selects suitable single-skill and multi-skill configurations, synthesizes tasks through rule-based and agent-based verification with feedback-guided repair, and retains only successful trajectories that substantially use every required skill. Using 2,000 public skills, SKT produces 4,000 task packages and 27,164 verified trajectories. Based on the same pipeline and a disjoint test pool, we further construct SkillEval, a held-out executable benchmark for evaluating skill use. Experiments across diverse models, benchmarks, and agent harnesses show that supervised fine-tuning on SKT-generated trajectories consistently improves skill-use performance. Verification ablations, cross-harness evaluation, and scaling experiments further demonstrate that these gains depend on high-quality supervision, extend beyond a single agent interface, and increase with broader skill coverage. Together, these results establish verified data synthesis as an effective and scalable approach for skill-use training.