컨텍스트가 물 때: 문서 수준 어텐션 붕괴를 통한 RAG 오염 탐지
When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse
August 7, 2026
저자: Yingtao Ren, Ziyi Zhao, Yiwei Fu, Xiao Luo, Yu-Cheng Chang, Chin-Teng Lin
cs.AI
초록
검색 증강 생성(RAG)은 대규모 언어 모델을 향상시키는 데 필수적이다. 그러나 RAG는 점점 더 중독 공격에 취약해지고 있으며, 이러한 공격에서는 적대적 문서가 주입되어 생성기 출력을 조작한다. 기존 방법들은 혼란도(perplexity) 및 일관성 검사와 같은 출력 측 신호에 의존하여 이러한 공격을 탐지한다. 그럼에도 불구하고, 우리의 분석은 의도적인 공격이 종종 거짓 확신을 유도하여, 중독된 출력이 정상 출력보다 오히려 더 낮은 혼란도를 보임으로써 불확실성 기반 탐지를 무력화한다는 것을 밝힌다. 이 문제를 해결하기 위해, 우리는 생성기의 내부 역학을 탐구하고 주의 집중 붕괴(Attention Collapse)라고 불리는 독특한 시그니처를 식별한다. 정상 생성에서의 분산된 주의와 달리, 공격받은 생성은 주의가 중독된 문서에 집중됨에 따라 엔트로피가 감소하는 것을 보인다. 이러한 발견을 바탕으로, 우리는 공격받은 생성을 식별하기 위해 주의 역학을 모니터링하는 경량 탐지 프레임워크인 D-SCAN(문서 수준 신호 붕괴 분석, Document-level Signal Collapse Analysis)을 제안한다. 여러 공격 벤치마크에 대한 광범위한 실험은 우리 방법의 효과성을 입증한다. 더욱이, D-SCAN은 공격이 최종 답변을 변경하지 못하는 경우에도 공격을 탐지할 수 있다. 코드는 https://github.com/yingtaoren/D-Scan.git에서 확인할 수 있다.
English
Retrieval-augmented generation (RAG) is indispensable for enhancing large language models. However, RAGs are increasingly susceptible to poisoning attacks, in which adversarial documents are injected to manipulate generator outputs. Previous methods rely on output-side signals such as perplexity and consistency checks to detect such attacks. Nevertheless, our analysis reveals that deliberate attacks often induce false confidence, where poisoned outputs exhibit even lower perplexity than benign ones, rendering uncertainty-based detection ineffective. To address this challenge, we explore the internal dynamics of the generator and identify a distinctive signature termed Attention Collapse. Unlike the dispersed attention in benign generations, attacked generations exhibit a decrease in entropy as attention concentrates on poisoned documents. Building on these findings, we propose D-SCAN (Document-level Signal Collapse Analysis), a lightweight detection framework that monitors attention dynamics to identify attacked generations. Extensive experiments on multiple attack benchmarks demonstrate the effectiveness of our method. Moreover, D-SCAN can detect attacks even when they fail to alter the final answer. Code is available at https://github.com/yingtaoren/D-Scan.git.