복사 가능한 컨텍스트에 기반한 안전장치는 LLM에 신뢰할 수 있는 안전성을 제공할 수 없다
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
July 30, 2026
저자: Pingyu Wu, Lingyao Zhu, Weiming Zhang, Nenghai Yu
cs.AI
초록
대규모 언어 모델의 안전장치는 답변이 어떻게 사용될지 확인하기 전에 응답 여부를 결정한다. 이는 이중 용도 과제에서 근본적인 문제를 야기한다. 동일한 답변이 인가된 전문가에게는 도움이 될 수 있지만 공격자에게도 도움이 될 수 있으며, 공격자는 정상적인 요청과 상호작용 이력을 모방할 수 있기 때문이다. 우리는 모델이 제공하는 역량과 하류 사용에 관한 이용 가능한 증거를 분리한다. 해당 증거가 복제 가능할 때, 유용한 답변을 보존하면서 공격자 지원의 정확한 최악의 하한을 도출한다. 그 결과는 안전 트릴레마를 산출한다. 즉, 유용한 역량, 신뢰할 수 있는 안전, 개방적 접근은 공존할 수 없다. 그런 다음 신뢰할 수 있는 자격 증명이 실제 하류 사용을 예측하는 복제하기 어려운 정보를 추가함으로써 기존 안전장치를 보완할 수 있는 방법을 보여주고, 하한을 제거하는 데 필요한 더 강력한 조건을 식별한다. 이중 용도 평가, 적대적 공격, 배포된 신뢰 접근 프로그램의 증거는 이러한 조건들의 실질적 관련성을 뒷받침한다.
English
Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and interaction history. We separate the capability released by the model from the evidence available about downstream use. When that evidence is copyable, we derive the exact worst-case floor on attacker assistance while preserving useful answers. The result yields a safety trilemma: Useful Capability, Reliable Safety, and Open Access cannot coexist. We then show how a trusted credential can complement existing safeguards by adding hard-to-copy information that predicts actual downstream use, and identify the stronger condition needed to eliminate the floor. Evidence from dual-use evaluations, adaptive attacks, and deployed trusted-access programs supports the practical relevance of these conditions.