基于可复制上下文的防护措施无法为大语言模型提供可靠的安全保障
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
July 30, 2026
作者: Pingyu Wu, Lingyao Zhu, Weiming Zhang, Nenghai Yu
cs.AI
摘要
大语言模型的安全防护措施在尚未得知答案将如何被使用之前,就已决定是否作答。这给双重用途任务带来了一个基本问题:同一个答案既可能帮助授权专业人员,也可能帮助攻击者,而攻击者能够模仿良性请求和交互历史。我们将模型释放的能力与关于下游使用的可用证据分离开来。当这些证据可复制时,我们推导出了在保留有用回答的同时,攻击者所获协助的确切最坏情况下限。这一结果导致了安全三难困境:有用能力、可靠安全与开放访问无法共存。我们随后展示了可信凭证如何通过增添难以复制的信息来补充现有防护措施,这些信息能够预测实际的下游使用情况;我们还指出了消除该下限所需的更强条件。来自双重用途评估、自适应攻击和已部署的可信访问计划的证据,支持了这些条件的实际相关性。
English
Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and interaction history. We separate the capability released by the model from the evidence available about downstream use. When that evidence is copyable, we derive the exact worst-case floor on attacker assistance while preserving useful answers. The result yields a safety trilemma: Useful Capability, Reliable Safety, and Open Access cannot coexist. We then show how a trusted credential can complement existing safeguards by adding hard-to-copy information that predicts actual downstream use, and identify the stronger condition needed to eliminate the floor. Evidence from dual-use evaluations, adaptive attacks, and deployed trusted-access programs supports the practical relevance of these conditions.