ChatPaper.aiChatPaper

從重新加權到重寫:解鎖具影響力樣本在訓練資料歸因中的干預效應

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

September 2, 2026
作者: Yuzhang Luo, Chenpeng Wang, Jianhui Chen, Liangming Pan
cs.AI

摘要

訓練資料歸因(TDA)旨在識別塑造模型行為的訓練範例,但其干預價值取決於選取了哪些範例以及如何修改這些範例。影響函數(IF)估計在無窮小重新加權下的行為變化,然而在傳統基於權重的干預下,IF 選出的範例往往相較於隨機選取僅展現有限優勢。這引出一個問題:具影響力的範例是否缺乏干預價值,還是重新加權未能實現其行為槓桿。我們提出影響引導的回應重寫,利用 IF 識別干預目標,並將其回應替換為行為一致或行為相反的監督,同時保持指令固定。在四個開放權重的大型語言模型中,我們以知識性棄答作為主要實驗場域,對相同的影響函數選出之範例,比較重寫與重新加權。回應重寫產生更強、更持久且雙向的行為轉變,而對相同範例進行重新加權則產生微弱且不一致的效果。進一步分析顯示,影響選出的範例比替代選擇器提供更大的重寫槓桿,且變化仍集中於與目標相關的行為。相同的質性對比也延伸到安全拒絕。這些結果將影響估計所捕捉的局部重新加權效應,與其所識別範例更廣泛的干預槓桿區分開來,促使對 TDA 方法進行干預感知評估。
English
Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweighting, yet IF-selected examples often show limited advantages over random selection under conventional weight-based interventions. This raises the question of whether influential examples lack intervention value or whether reweighting fails to realize their behavioral leverage.We introduce influence-guided response rewriting, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed. Across four open-weight LLMs, we compare rewriting and reweighting on the same influence-selected examples using epistemic abstention as our primary testbed. Response rewriting produces stronger, more persistent, and bidirectional behavioral shifts, while reweighting the same examples yields weak and inconsistent effects. Further analyses show that influence-selected examples provide greater rewriting leverage than alternative selectors, with changes remaining concentrated on target-relevant behaviors. The same qualitative contrast extends to safety refusal. These results distinguish the local reweighting effects captured by influence estimates from the broader intervention leverage of the examples they identify, motivating intervention-aware evaluation of TDA methods.