ChatPaper.aiChatPaper

DiSCO:通过分布引导的对比提示优化防御文本到图像生成

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

August 17, 2026
作者: Tong Zhang, Motasem Alfarra, Carlos Hinojosa, Christos Louizos, Bernard Ghanem
cs.AI

摘要

随着文本到图像生成模型的进步,它们引发了严重的安全问题,尤其是生成暴力、裸露等不宜工作场所(NSFW)内容,而红队对抗性攻击进一步加剧了这一隐患。现有防御手段主要在白盒假设下运行,依赖文本编码器优化、权重编辑或推理时干预,从根本上无法扩展到专有模型。基于大语言模型提示词重写的黑盒替代方案具有更广泛的适用性,但在我们识别为良性对抗问题的关键场景中失效:即提示词在语言层面是安全的,但由于模型习得的数据分布仍会触发有害生成。我们提出DiSCO,一种零样本、严格黑盒的防御方法,完全在提示词层面作为即插即用模块运行,无需模型重训练、微调或访问模型内部信息。DiSCO通过束搜索执行分布引导的后缀扩展,并利用目标模型自身生成的安全与不安全图像池进行对比评分来优化,通过迭代自适应反馈直至生成安全内容。我们证明,在I2P基准上面对多种红队攻击时,DiSCO能够持续提升未防御和已防御模型的安全性,分别实现37.7%和25.13%的攻击成功率(ASR)降低,同时保持语义保真度并改善图像连贯性。作为黑盒、架构无关的模块,DiSCO可便捷地应用于任何文本到图像系统,无需对模型本身做任何改动。
English
As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or inference-time intervention, and fundamentally cannot scale to proprietary models. Black-box alternatives based on LLM prompt rewriting offer broader applicability, yet fail in a critical regime we identify as the benign adversarial problem: prompts that are linguistically safe but still trigger harmful generation due to the model's learned data distribution. We propose DiSCO, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals. DiSCO performs distribution-guided suffix expansion via beam search, optimized through contrastive scoring over safe and unsafe image pools generated by the target model itself, with iterative adaptive feedback until safe content is produced. We demonstrate that DiSCO consistently enhances the safety of both undefended and defended models on the I2P benchmark under multiple red-teaming attacks, achieving 37.7% and 25.13% ASR reduction, respectively, while maintaining semantic fidelity and improving image coherence. As a black-box, architecture-agnostic module, DiSCO can be readily applied to any text-to-image system without necessitating any changes to the model itself.