ChatPaper.aiChatPaper

SkillEvo: マルチターン対話フィードバックからの自己更新型進化勾配

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

August 13, 2026
著者: Qianxi Yan, Chunrong Chen, Jiuzhou Zhao, Min Zhang, Yongzhou Xu, Xiaochuan Xu
cs.AI

要旨

現在、エージェントスキルは手作業で作成されるか、単一のLLM生成パスで生成されるかのいずれかであり、その結果として、自身が実際に引き起こすインタラクションの失敗から改善するための閉ループを持たない。最近の研究ではこのループを閉じているが、そのフィードバックは単一ターンの質問応答評価から得られている。その結果として生じるのは著しい非対称性である。すなわち、最初のラウンドが単一のやり取りで明らかになるギャップを修正してしまうと、進化の勾配は減衰し、複数ターンにわたって初めて表面化する欠陥は不可視のままとなり、進化は停滞する。これらのシステムにおけるガバナンスも同様に、エンドツーエンドの検証スコアによって駆動される。これはスカラーゲートであり、劣化した候補を拒否することはできるが、その構造的原因を特定することも修復することもできない。我々は、持続的なスキル進化における制約要因は、編集能力でも反復回数でもなく、評価フィードバックが信頼できる進化勾配を供給し続けるかどうかであると主張する。我々はSkillEvoを提案する。SkillEvoでは、信頼できるフィードバックが勾配を生成し、制御可能なガバナンスがその方向を制約する。最初の構成要素は、マルチターンユーザーシミュレーションを、評価エンドポイントからフィードバック生成器へと転換する。すなわち、フォローアップ質問が欠陥を層ごとに露呈することで、改訂の各ラウンドがフィードバックを消費すると同時に新しいフィードバックを生成する。第二の構成要素は、スカラーゲートによる受動的な拒否を、事実性の劣化と構造的肥大を能動的に修復する独立したガバナンス層に置き換え、劣化が蓄積するにつれて勾配が逸脱するのを防ぐ。クラウドサービスの6カテゴリ、9つの本番スキル、98のスキル参照ファイルにわたって、SkillEvoは自己内省ベースの進化を23.0ポイント、単一ターンQA駆動の進化を15.4ポイント上回る。
English
Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls. Governance in these systems is likewise driven by an end-to-end verification score, a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause. We argue that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction. The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback. The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates. Across six categories of cloud services, 9 production Skills, and 98 skill-reference files, SkillEvo surpasses self-reflection-based evolution by 23.0 points and single- turn-QA-driven evolution by 15.4 points.