以善為目的之對抗性攻擊:貫穿視覺內容生命週期的主動防護綜述
Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle
August 5, 2026
作者: Jiaming Zhang, Boyang Chen, Zherui Li, Fuyao Zhang, Xinyu Yan, Hong Xi Tae, Wenwen He, Xuan Wang, Siqi Guo, Junhao Dong, Kun Wang, Hanxun Huang, Yige Li, Xingjun Ma, Yang Cao, Lingjuan Lyu, Wei Yang Bryan Lim
cs.AI
摘要
一旦視覺內容進入AI管道,其擁有者往往對其使用方式幾乎無法保留技術上的控制權。法律與監管救濟措施可以處理濫用問題,但許多技術性干預必須更早應用,即在內容發布或存取之時。本综述檢視了圍繞此干預點所形成的防護範式,我們稱之為「為善而用的對抗性攻擊」。長期以來被研究為對學習模型之攻擊的擾動與結構化訊號,如今反而由資料擁有者、創作者、平台或稽核人員所應用,用以破壞未經授權的自動化處理或支援後續的課責機制。五個研究社群基本上各自獨立地達成了此一逆向應用,各自因應視覺資產生命週期的不同階段:分享時的隱私濾波器以對抗非預期的辨識、不可學習範例以對抗未經授權的訓練、生成式防護機制以對抗惡意編輯或模仿、用於存取控制的對抗性CAPTCHA以對抗自動化代理程式,以及用於流通後歸屬追溯的來源驗證機制。儘管這些方法在不同的場域中發展,且成功標準互不相容,其中許多方法利用了人類知覺、語意詮釋與機器推論之間持續存在的落差,這顯示此範式在視覺管道朝向多模態模型與自主代理演進之際仍具有相關性。為了使各方法的主張具有可比性,我們沿著可遷移性、適應性與部署就緒程度等共同軸向來評估這五個家族。貫穿整個生命週期,我們發現大多數防護機制仍主要針對靜態或弱適應性對手進行驗證,而受控基準測試之外的證據仍然稀缺。我們最後整合了跨階段的對策與開放性問題,以實現穩健、可組合且可部署的所有者端防護。
English
Once visual content enters an AI pipeline, its owner often retains little technical control over how it is used. Legal and regulatory remedies can address misuse, but many technical interventions must be applied earlier, when content is released or accessed. This survey examines the protective paradigm that has grown around this intervention point, which we call adversarial attacks for good. Perturbations and structured signals long studied as attacks on learned models are instead applied by data owners, creators, platforms, or auditors to disrupt unauthorized automation or support later accountability. Five research communities have arrived at this inversion largely independently, each addressing a different stage of a visual asset's lifecycle: privacy filters against unwanted recognition at sharing time, unlearnable examples against unauthorized training, generative safeguards against malicious editing or imitation, adversarial CAPTCHAs for access control against automated agents, and provenance mechanisms for post-circulation attribution. Although developed in separate venues with incompatible success criteria, many of these methods exploit persistent gaps between human perception, semantic interpretation, and machine inference, suggesting that the paradigm remains relevant as visual pipelines evolve toward multimodal models and autonomous agents. To make their claims comparable, we evaluate all five families along shared axes of transferability, adaptability, and deployment readiness. Across the lifecycle, we find that most protections are still validated mainly against static or weakly adaptive adversaries, while evidence beyond controlled benchmarks remains scarce. We close by consolidating cross-stage countermeasures and open problems for robust, composable, and deployable owner-side protection.