재가중에서 재작성으로: 훈련 데이터 기여도에서 영향력 있는 샘플의 개입 효과 밝히기
From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
September 2, 2026
저자: Yuzhang Luo, Chenpeng Wang, Jianhui Chen, Liangming Pan
cs.AI
초록
훈련 데이터 기여도(TDA)는 모델 행동을 형성하는 훈련 예제를 식별하는 것을 목표로 하지만, 그 개입 가치는 어떤 예제가 선택되는지와 그것들이 어떻게 수정되는지 모두에 달려 있다. 영향 함수(IF)는 무한소 재가중 하에서의 행동 변화를 추정하지만, IF로 선택된 예제들은 기존의 가중치 기반 개입 하에서 무작위 선택 대비 제한적인 이점만을 보이는 경우가 많다. 이는 영향력 있는 예제들에 개입 가치가 없는 것인지, 아니면 재가중이 그들의 행동적 지렛대를 실현하지 못하는 것인지라는 질문을 제기한다. 우리는 영향 함수 기반 응답 재작성(influence-guided response rewriting)을 도입하는데, 이는 IF를 사용하여 개입 대상을 식별하고 지시문은 고정한 채 그 응답을 행동 정렬형 또는 행동 반대형 지도 신호로 대체한다. 네 개의 오픈 웨이트 LLM 전반에 걸쳐, 우리는 인식적 유보를 주요 테스트베드로 삼아 동일한 영향 선택 예제들에 대한 재작성과 재가중을 비교한다. 응답 재작성은 더 강력하고 더 지속적이며 양방향성의 행동 변화를 만들어내는 반면, 동일한 예제에 대한 재가중은 약하고 일관되지 않은 효과를 낸다. 추가 분석은 영향 선택 예제들이 대체 선택자들보다 더 큰 재작성 지렛대를 제공하며, 변화가 목표 관련 행동에 집중된 채 남아 있음을 보여준다. 동일한 질적 대비는 안전 거부로도 확장된다. 이러한 결과는 영향 추정값이 포착하는 국소적 재가중 효과와 그것들이 식별하는 예제들의 더 광범위한 개입 지렛대를 구별하며, TDA 방법의 개입 인지 평가에 동기를 부여한다.
English
Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweighting, yet IF-selected examples often show limited advantages over random selection under conventional weight-based interventions. This raises the question of whether influential examples lack intervention value or whether reweighting fails to realize their behavioral leverage.We introduce influence-guided response rewriting, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed. Across four open-weight LLMs, we compare rewriting and reweighting on the same influence-selected examples using epistemic abstention as our primary testbed. Response rewriting produces stronger, more persistent, and bidirectional behavioral shifts, while reweighting the same examples yields weak and inconsistent effects. Further analyses show that influence-selected examples provide greater rewriting leverage than alternative selectors, with changes remaining concentrated on target-relevant behaviors. The same qualitative contrast extends to safety refusal. These results distinguish the local reweighting effects captured by influence estimates from the broader intervention leverage of the examples they identify, motivating intervention-aware evaluation of TDA methods.