ChatPaper.aiChatPaper

基於可複製上下文的保護措施無法為大型語言模型提供可靠的安全性

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs

July 30, 2026
作者: Pingyu Wu, Lingyao Zhu, Weiming Zhang, Nenghai Yu
cs.AI

摘要

大型語言模型的防護措施在決定是否回答時,尚未看到答案將如何被使用。這對雙重用途任務構成一個基本問題:相同的答案既可能幫助授權專業人士,也可能幫助攻擊者,而攻擊者可以模仿良性的請求與互動歷史。我們將模型所釋放的能力,與關於下游使用的可得證據區分開來。當該證據是可複製的,我們推導出在保留有用答案的同時,攻擊者所能獲得協助的確切最壞情況下限。此結果產生一個安全三難困境:有用能力、可靠安全與開放取用無法共存。接著我們展示,可信憑證如何透過加入難以複製且能預測實際下游使用的資訊,來補充現有防護措施,並找出消除該下限所需的更嚴格條件。來自雙重用途評估、適應性攻擊及已部署的可信存取計畫的證據,支持了這些條件的實際相關性。
English
Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and interaction history. We separate the capability released by the model from the evidence available about downstream use. When that evidence is copyable, we derive the exact worst-case floor on attacker assistance while preserving useful answers. The result yields a safety trilemma: Useful Capability, Reliable Safety, and Open Access cannot coexist. We then show how a trusted credential can complement existing safeguards by adding hard-to-copy information that predicts actual downstream use, and identify the stronger condition needed to eliminate the floor. Evidence from dual-use evaluations, adaptive attacks, and deployed trusted-access programs supports the practical relevance of these conditions.