교사가 오도할 때: 허위 신호를 인식하는 온-폴리시 증류
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
August 4, 2026
저자: Yinuo Jiang, Yongjie Ye, Zhou Tao, Xiang Zhuang, Qiang Zhang, Huajun Chen, Tiankai Li
cs.AI
초록
온-폴리시 증류(OPD)는 학생 모델이 샘플링한 궤적에 대해 토큰 수준의 밀집 교사 신호를 제공함으로써 교사 모델의 능력을 전이한다. 최근의 선택적 OPD 방법들은 신뢰도가 높거나, 정보량이 풍부하거나, 학습 가능한 신호를 우선시함으로써 이 과정을 개선한다. 그러나 이러한 가정들은 언어 모델의 근본적인 실패 양상을 간과한다. 즉, 토큰 수준 판단이 작업 특정 증거보다는 입력과 무관한 언어 사전 정보, 형식 관행, 또는 고정관념적 추론 템플릿에 의해 주도될 수 있다는 점이다. 우리는 이러한 최적화 관련성을 지니지만 입력 근거가 약한 감독 신호를 OPD에서 허위 신호(spurious signal)라고 지칭하며, 이는 작업 개선 방향에 거의 기여하지 못하면서 큰 기울기를 생성할 수 있다. 이 문제를 완화하기 위해 우리는 입력 근거성과 최적화 영향을 기반으로 오도된 토큰 수준 감독을 식별하고 필터링하는 SA-OPD(허위 신호 인지 온-폴리시 증류 프레임워크)를 제안한다. SA-OPD는 토큰 수준 증류 신호가 입력에 실제로 의존하는지를 추정하는 경량의 입력 근거성 프록시를 도입한다. 그런 다음 낮은 입력 근거성과 극단적인 증류 발산을 동시에 나타내는 토큰만을 필터링함으로써, 영향력이 큰 허위 업데이트를 제거하고 세밀한 수준의 OPD 최적화를 달성한다. 대규모 언어 모델(LLM) 및 비전-언어 모델(VLM) 환경 모두에 대한 광범위한 실험은 SA-OPD가 기본 OPD(Vanilla OPD) 및 경쟁력 있는 선택적 방법들을 일관되게 능가함을 보여준다. 이러한 결과는 입력 근거성을 OPD 감독 선택의 핵심 차원으로 확립하며, 허위 업데이트 완화를 위한 간단하고 효과적인 전략을 제공한다.
English
On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable. However, the assumptions overlook a fundamental failure mode of language models: their token-level judgments can be driven by input-agnostic language priors, formatting conventions, or stereotyped reasoning templates rather than task-specific evidence. We refer to such optimization-relevant but weakly input-grounded supervision as spurious signals in OPD, which may produce large gradients while contributing little task-improving direction. To mitigate this issue, we propose SA-OPD, a Spurious-Signal-Aware On-Policy Distillation framework that identifies and filters misleading token-level supervision based on input-groundedness and optimization impact. SA-OPD introduces a lightweight input-groundedness proxy estimating whether a token-level distillation signal truly depends on the input. It then filters only tokens that simultaneously exhibit low input-groundedness and extreme distillation divergence, thereby removing high-impact spurious updates and achieving fine-grained OPD optimization. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that SA-OPD consistently outperforms Vanilla OPD and competitive selective methods. These results establish input-groundedness as a key dimension for OPD supervision selection and offer a simple, effective strategy for mitigating spurious updates.