ChatPaper.aiChatPaper

再重み付けから書き換えへ:訓練データアトリビューションにおける影響力のあるサンプルの介入効果を解き放つ

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

September 2, 2026
著者: Yuzhang Luo, Chenpeng Wang, Jianhui Chen, Liangming Pan
cs.AI

要旨

訓練データ帰属(TDA)は、モデルの振る舞いを形成する訓練事例を特定することを目的とするが、その介入価値は、どの事例が選択されるかと、それらがどのように修正されるかの両方に依存する。影響関数(IF)は、微小な重み付け変更の下での行動変化を推定するが、IFによって選択された事例は、従来の重みベースの介入の下では、ランダム選択に対する限られた優位性しか示さないことが多い。このことは、影響力のある事例が介入価値を欠いているのか、あるいは重み付け変更がそれらの行動的レバレッジを実現できていないのかという疑問を提起する。我々は、影響誘導応答書き換えを導入する。これは、IFを用いて介入対象を特定し、指示を固定したまま、それらの応答を行動整合的または行動対立的な教師信号に置き換える。4つのオープンウェイトLLMにわたり、我々は、認識論的棄権を主要なテストベッドとして用いて、同じ影響選択された事例に対して書き換えと重み付け変更を比較する。応答書き換えは、より強力で、より持続的で、双方向の行動変化を生み出すが、同じ事例の重み付け変更は弱く一貫性のない効果しかもたらさない。さらなる分析は、影響選択された事例が、代替の選択手法よりも大きな書き換えレバレッジを提供し、変化がターゲット関連行動に集中したままであることを示す。同じ定性的な対比は、安全拒否にも拡張される。これらの結果は、影響推定によって捉えられた局所的な重み付け変更効果と、それらが特定する事例のより広範な介入レバレッジとを区別し、TDA手法の介入を考慮した評価を動機付ける。
English
Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweighting, yet IF-selected examples often show limited advantages over random selection under conventional weight-based interventions. This raises the question of whether influential examples lack intervention value or whether reweighting fails to realize their behavioral leverage.We introduce influence-guided response rewriting, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed. Across four open-weight LLMs, we compare rewriting and reweighting on the same influence-selected examples using epistemic abstention as our primary testbed. Response rewriting produces stronger, more persistent, and bidirectional behavioral shifts, while reweighting the same examples yields weak and inconsistent effects. Further analyses show that influence-selected examples provide greater rewriting leverage than alternative selectors, with changes remaining concentrated on target-relevant behaviors. The same qualitative contrast extends to safety refusal. These results distinguish the local reweighting effects captured by influence estimates from the broader intervention leverage of the examples they identify, motivating intervention-aware evaluation of TDA methods.