ChatPaper.aiChatPaper

MANCE: 매니폴드 인식 개념 소거

MANCE: Manifold Aware Concept Erasure

July 4, 2026
저자: Matan Avitan, Yoav Goldberg, Yanai Elazar
cs.AI

초록

개념 제거는 표현(representation)에서 특정 대상 개념을 제거하면서도 표현에 인코딩된 다른 정보들은 보존하는 것을 목표로 한다. 이는 표현이 종종 제거 대상 개념과 상관관계가 있는 많은 개념들을 인코딩하기 때문에 어려운 작업이며, 따라서 대상 개념을 제거하면 다른 정보들이 손상될 위험이 있다. 우리는 다양체 제약 가설(Manifold Constraint Hypothesis, MCH)을 제안한다: 만약 자연적 표현들이 구조화된 저차원 다양체에 집중된다면, 중재(intervention)는 해당 다양체로 제약되어야 하며 중재 과정에서 표현에 인코딩된 다른 정보를 더 잘 보존할 수 있다는 것이다. 우리는 MCH를 새로운 개념 제거 방법인 MANCE(Manifold aware Concept Erasure)에 구현한다. MANCE는 대상 개념을 예측하는 분류기에서 얻은 신호를 사용하여 표현에 대한 반복적 업데이트를 수행한다. 자연 입력으로부터 얻은 표현을 사용해 다양체를 추정한 후, 개념 제거 업데이트를 추정된 다양체로 투영한다. 우리는 텍스트와 비전을 아우르는 119개 설정(13개 언어 모델, 3개 NLP 개념, 40개 CelebA-CLIP 속성 포함)에 대해 광범위한 평가를 수행했다. 기존 방법들 위에 MANCE를 적용한 결과, 누출(leakage) 결과가 일관되게 개선되었다. 또한 MANCE+와 MANCE++를 도입했는데, 이들은 MANCE를 적용하기 전에 닫힌 형태 제거 알고리즘을 먼저 수행함으로써, 동일한 전체 공간 업데이트와 비교하여 더 나은 누출-정밀성(surgicality) 상충 관계를 달성한다. 우리의 최고 방법인 MANCE++는 비선형 개념 제거에서 최첨단 결과를 달성한다. 이러한 결과는 제거 설정에서 MCH를 지지한다: 중재는 자연적 표현 다양체로 제약되어야 한다.
English
Concept erasure aims to remove a target concept from a representation while preserving the other information encoded in it. This is difficult because representations encode many concepts that are often correlated with the erasure target, so removing the target risks damaging them. We propose the Manifold Constraint Hypothesis (MCH): if natural representations concentrate on a structured, lower-dimensional manifold, then interventions should be constrained to that manifold and better preserve other information encoded in the representation during interventions. We instantiate MCH in a new concept erasure method: MANifold aware Concept Erasure (MANCE). MANCE performs iterative updates to the representations using signals from a classifier that predicts a target concept. We estimate the manifold using representations obtained from natural inputs, and then we project the concept removal update to the estimated manifold. We perform extensive evaluation on 119 settings spanning text and vision, including 13 language models, three NLP concepts, and 40 CelebA-CLIP attributes. Employing MANCE on top of previous methods shows consistent improved leakage results. We also introduce MANCE+ and MANCE++, which prepend a closed-form erasure algorithm before employing MANCE, achieving better leakage--surgicality tradeoffs relative to matched full-space updates. MANCE++, our best method, achieves state-of-the-art results on nonlinear concept erasure. These results support MCH in the erasure setting: interventions should be constrained to the natural representation manifold.