인기가 많을수록 잊기 어렵다: LLM 언러닝을 위한 적응형 인기도
The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning
August 14, 2026
저자: Anna Borisiuk, Andrey Savchenko, Alexander Panchenko, Elena Tutubalina
cs.AI
초록
인기 있는 사실은 사전 학습 중 더 깊이 기억되며 희귀한 사실보다 제거에 더 오래 저항하지만, 기존 LLM 언러닝 방법은 학습 데이터 빈도와 관계없이 균일한 그래디언트 압력을 적용한다. 우리는 AdaPop(Adaptive Popularity) 방법을 제안한다. 이 방법은 로컬 토큰 신뢰도와 외부 프록시(예: Wikidata 사이트링크, LLM-as-Judge)에서 파생된 사실별 인기도 의존 지수를 결합하고, 에폭마다 유지(retain) 손실 페널티를 조정하는 이중 상승 제어기를 통해 망각-유지 균형을 자동화한다. 세 가지 모델 계열과 두 가지 벤치마크에서 AdaPop은 의역된 쿼리에서 경쟁 방법보다 망각된 콘텐츠를 약 5배 적게 누출했고, 적대적 재구성에서는 약 1.6배 적게 누출했다. 우리는 내부 지표로 분석을 뒷받침한다. 우리 방법에서는 망각 집합의 은닉 상태가 언러닝 전 모델의 상태로부터 다른 방법보다 더 멀리 이동하는 반면, 유지 집합의 표현은 가깝게 유지된다.
English
Popular facts are memorised more deeply during pretraining and resist removal longer than rare ones, yet existing LLM unlearning methods apply uniform gradient pressure regardless of training-data frequency. We propose the AdaPop (Adaptive Popularity) method, which combines local token confidence with a per-fact popularity-dependent exponent derived from an external proxy (e.g., Wikidata sitelinks, LLM-as-Judge), and automates the forget-retain balance via a dual-ascent controller that adjusts the retain penalty each epoch. Across three model families and two benchmarks, AdaPop leaks ~5x less forgotten content than competing methods under paraphrased queries and ~1.6x less under adversarial reformulations. We support our analysis with internal metrics: under our method, forget-set hidden states move further from the pre-unlearning model's states than under other methods, while retain-set representations remain close.