複製可能なコンテキストに基づく安全策は、大規模言語モデルに信頼できる安全性を提供することはできない
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
July 30, 2026
著者: Pingyu Wu, Lingyao Zhu, Weiming Zhang, Nenghai Yu
cs.AI
要旨
大規模言語モデルの安全対策は、回答がどのように使用されるかを確認する前に回答の可否を判断する。これは二重用途タスクにおいて根本的な問題を生じさせる。すなわち、同じ回答が正規の専門家にも攻撃者にも役立ち得る一方で、攻撃者は正規のリクエストと対話履歴を模倣することができるからである。我々は、モデルが解放する能力と、下流での使用に関して利用可能な証拠とを分離する。その証拠が複製可能である場合、有用な回答を維持しつつ、攻撃者支援の正確な最悪ケースの下限を導出する。この結果は安全上の三律背反をもたらす。すなわち、有用な能力、信頼できる安全性、オープンアクセスの三者は同時には成立し得ない。さらに我々は、実際の下流での使用を予測する複製困難な情報を付加することにより、信頼できる認証情報が既存の安全対策を補完できることを示し、その下限を排除するために必要となるより強い条件を特定する。二重用途評価、適応的攻撃、および展開済みの信頼アクセスプログラムからの証拠は、これらの条件の実践的妥当性を支持する。
English
Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and interaction history. We separate the capability released by the model from the evidence available about downstream use. When that evidence is copyable, we derive the exact worst-case floor on attacker assistance while preserving useful answers. The result yields a safety trilemma: Useful Capability, Reliable Safety, and Open Access cannot coexist. We then show how a trusted credential can complement existing safeguards by adding hard-to-copy information that predicts actual downstream use, and identify the stronger condition needed to eliminate the floor. Evidence from dual-use evaluations, adaptive attacks, and deployed trusted-access programs supports the practical relevance of these conditions.