UNMASK: 텍스트 분류기에서 허위 지름길 발견 및 인과적 검증
UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers
August 10, 2026
저자: Chidaksh Ravuru, Shashank Srivastava
cs.AI
초록
대규모 크라우드소싱 말뭉치로 훈련된 신경 언어 모델은 진정한 언어적 또는 인과적 관련성 없이 목표 레이블에 결합된 허위 표면 패턴을 자주 이용하여, 적대적 입력이나 분포 외 입력에서는 실패하면서도 벤치마크 성능을 끌어올린다. 기존 접근법은 특징 어휘의 수동 지정을 요구하거나 발견 과정을 부분적으로만 자동화하여, 데이터셋 수준 상관관계와 모델 수준 이용 사이의 격차를 해결하지 못한다. 우리는 추가적인 인간 주석 없이 텍스트 분류기의 허위 상관관계를 발견, 인과 검증, 완화하는 완전 자동화 파이프라인 U N M ASK를 제시한다. 라벨이 없는 훈련 예제가 주어지면 U N M ASK는 실행 가능한 부울 표현식으로 후보 표면 패턴을 생성하고, 독립적 반복을 포함한 통계적 검증 프로토콜로 이를 필터링하며, 검증된 반사실적 개입을 통해 인과적 모델 의존성을 확립한다. 이후 인과적으로 확인된 특징들은 Deep Feature Reweighting(DFR)을 위한 주석 없는 그룹 정의로 사용되어, 표준 DFR이 요구하는 그룹 레이블을 제거한다. MNLI로 훈련된 BERT와 RoBERTa에 적용한 결과, 우리의 파이프라인은 기존에 알려진 어휘 중첩 및 부정 편향을 독립적으로 재발견하여 BERT에서는 10개 중 9개, RoBERTa에서는 6개의 특징을 검증했고, HANS 정확도를 최대 12.58퍼센트 포인트 향상시켰다. CivilComments-WILDS에서는 프로그램 기반 그룹이 인구통계학적 주석 없이도 수동 레이블 DFR(Kirichenko et al., 2023)의 최악 그룹 정확도 70.1%와 일치하는 성능을 보였다. 또한 발견 및 검증 단계가 보상 모델 선호 데이터로 일반화됨을 입증하여, RewardBench2에서 해석 가능한 허위 상관관계를 표면화한다.
English
Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on adversarial or out-of-distribution inputs. Existing approaches either require manual specification of the feature vocabulary or automate discovery only partially, leaving the gap between dataset-level correlation and model-level exploitation unaddressed. We present U N M ASK, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation. Given unlabeled training examples, U N M ASK generates candidate surface patterns as executable boolean expressions, filters them through a statistical validation protocol with independent replication, and establishes causal model dependence via verified counterfactual interventions. Causally confirmed features then serve as annotation-free group definitions for Deep Feature Reweighting, eliminating the group labels that standard DFR requires. Applied to BERT and RoBERTa trained on MNLI, our pipeline independently rediscovers established lexical-overlap and negation biases, verifying 9 of 10 features on BERT and 6 on RoBERTa, and improving HANS accuracy by up to 12.58 pp. On CivilComments-WILDS, programmatic groups match the 70.1% worst- group accuracy of hand-labeled DFR (Kirichenko et al., 2023) without demographic annotation. We further demonstrate that the discovery and validation stages generalize to reward model preference data, surfacing interpretable spurious correlations in RewardBench2.