MANCE: Manifoldbewuste Conceptverwijdering
MANCE: Manifold Aware Concept Erasure
July 4, 2026
Auteurs: Matan Avitan, Yoav Goldberg, Yanai Elazar
cs.AI
Samenvatting
Conceptverwijdering heeft als doel een doelconcept uit een representatie te verwijderen terwijl de andere informatie die erin is gecodeerd, behouden blijft. Dit is moeilijk omdat representaties veel concepten coderen die vaak gecorreleerd zijn met het te verwijderen doel, waardoor het verwijderen van het doel risico loopt deze te beschadigen. We stellen de Manifold Constraint Hypothesis (MCH) voor: als natuurlijke representaties zich concentreren op een gestructureerd, lager-dimensionaal manifold, dan moeten interventies worden beperkt tot dat manifold en andere informatie die in de representatie is gecodeerd beter behouden tijdens interventies. We instantiëren MCH in een nieuwe methode voor conceptverwijdering: MANifold aware Concept Erasure (MANCE). MANCE voert iteratieve updates uit op de representaties met behulp van signalen van een classifier die een doelconcept voorspelt. We schatten het manifold met behulp van representaties verkregen uit natuurlijke invoer, en vervolgens projecteren we de update voor conceptverwijdering op het geschatte manifold. We voeren een uitgebreide evaluatie uit op 119 instellingen die tekst en visie omvatten, waaronder 13 taalmodellen, drie NLP-concepten en 40 CelebA-CLIP-attributen. Het toepassen van MANCE bovenop eerdere methoden toont consistent verbeterde lekkageresultaten. We introduceren ook MANCE+ en MANCE++, die een algoritme voor conceptverwijdering in gesloten vorm toevoegen vóór het gebruik van MANCE, waardoor een betere afweging tussen lekkage en chirurgische precisie wordt bereikt in vergelijking met overeenkomstige updates in de volledige ruimte. MANCE++, onze beste methode, behaalt state-of-the-art resultaten op niet-lineaire conceptverwijdering. Deze resultaten ondersteunen MCH in de verwijderingscontext: interventies moeten worden beperkt tot het natuurlijke representatiemanifold.
English
Concept erasure aims to remove a target concept from a representation while preserving the other information encoded in it. This is difficult because representations encode many concepts that are often correlated with the erasure target, so removing the target risks damaging them. We propose the Manifold Constraint Hypothesis (MCH): if natural representations concentrate on a structured, lower-dimensional manifold, then interventions should be constrained to that manifold and better preserve other information encoded in the representation during interventions. We instantiate MCH in a new concept erasure method: MANifold aware Concept Erasure (MANCE). MANCE performs iterative updates to the representations using signals from a classifier that predicts a target concept. We estimate the manifold using representations obtained from natural inputs, and then we project the concept removal update to the estimated manifold. We perform extensive evaluation on 119 settings spanning text and vision, including 13 language models, three NLP concepts, and 40 CelebA-CLIP attributes. Employing MANCE on top of previous methods shows consistent improved leakage results. We also introduce MANCE+ and MANCE++, which prepend a closed-form erasure algorithm before employing MANCE, achieving better leakage--surgicality tradeoffs relative to matched full-space updates. MANCE++, our best method, achieves state-of-the-art results on nonlinear concept erasure. These results support MCH in the erasure setting: interventions should be constrained to the natural representation manifold.