ChatPaper.aiChatPaper

善意对抗性攻击:视觉内容全生命周期的主动保护研究综述

Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle

August 5, 2026
作者: Jiaming Zhang, Boyang Chen, Zherui Li, Fuyao Zhang, Xinyu Yan, Hong Xi Tae, Wenwen He, Xuan Wang, Siqi Guo, Junhao Dong, Kun Wang, Hanxun Huang, Yige Li, Xingjun Ma, Yang Cao, Lingjuan Lyu, Wei Yang Bryan Lim
cs.AI

摘要

一旦视觉内容进入AI处理流程,其所有者往往难以在技术上控制内容的使用方式。法律和监管手段可以应对滥用行为,但许多技术干预必须在更早阶段进行,即在内容发布或访问之时。本综述考察了围绕该干预点发展起来的保护范式,我们称之为“为善的对抗攻击”。长期作为对已学习模型的攻击来研究的扰动和结构化信号,如今反而被数据所有者、创作者、平台或审计者用来破坏未经授权的自动化,或为事后问责提供支持。五个研究社区在很大程度上独立地实现了这一反转,各自针对视觉资产生命周期的不同阶段:共享时的隐私过滤器以抵御非预期的识别,不可学习样本以防御未经授权的训练,生成式防护以抵御恶意编辑或模仿,对抗性验证码以对自动化代理进行访问控制,以及溯源机制以支持传播后的归因。尽管这些方法在不同的场合中提出,评价标准也互不兼容,但其中许多方法利用了人类感知、语义理解与机器推理之间持续存在的差距,这表明随着视觉处理流程向多模态模型和自主代理演进,该范式仍然具有相关性。为使各方结论可比,我们沿着可迁移性、适应性和部署就绪度三个共同维度评估这五个方法家族。纵观整个生命周期,我们发现大多数保护措施仍主要针对静态或弱自适应的对手进行验证,而超出受控基准测试的证据仍然稀缺。最后,我们整合了跨阶段的对策与开放问题,以推动健壮、可组合且可部署的所有者侧保护。
English
Once visual content enters an AI pipeline, its owner often retains little technical control over how it is used. Legal and regulatory remedies can address misuse, but many technical interventions must be applied earlier, when content is released or accessed. This survey examines the protective paradigm that has grown around this intervention point, which we call adversarial attacks for good. Perturbations and structured signals long studied as attacks on learned models are instead applied by data owners, creators, platforms, or auditors to disrupt unauthorized automation or support later accountability. Five research communities have arrived at this inversion largely independently, each addressing a different stage of a visual asset's lifecycle: privacy filters against unwanted recognition at sharing time, unlearnable examples against unauthorized training, generative safeguards against malicious editing or imitation, adversarial CAPTCHAs for access control against automated agents, and provenance mechanisms for post-circulation attribution. Although developed in separate venues with incompatible success criteria, many of these methods exploit persistent gaps between human perception, semantic interpretation, and machine inference, suggesting that the paradigm remains relevant as visual pipelines evolve toward multimodal models and autonomous agents. To make their claims comparable, we evaluate all five families along shared axes of transferability, adaptability, and deployment readiness. Across the lifecycle, we find that most protections are still validated mainly against static or weakly adaptive adversaries, while evidence beyond controlled benchmarks remains scarce. We close by consolidating cross-stage countermeasures and open problems for robust, composable, and deployable owner-side protection.