ChatPaper.aiChatPaper

善用のための敵対的攻撃:視覚コンテンツのライフサイクル全体にわたる能動的保護のサーベイ

Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle

August 5, 2026
著者: Jiaming Zhang, Boyang Chen, Zherui Li, Fuyao Zhang, Xinyu Yan, Hong Xi Tae, Wenwen He, Xuan Wang, Siqi Guo, Junhao Dong, Kun Wang, Hanxun Huang, Yige Li, Xingjun Ma, Yang Cao, Lingjuan Lyu, Wei Yang Bryan Lim
cs.AI

要旨

視覚コンテンツがAIパイプラインに入ると、その所有者は、コンテンツがどのように使用されるかについて技術的な制御をほとんど維持できないことが多い。法的・規制上の救済手段は悪用に対処できるが、多くの技術的介入は、コンテンツが公開またはアクセスされる時点という、より早い段階で適用されなければならない。本サーベイは、この介入点を中心に発展してきた保護的パラダイムを考察する。我々はこれを「善のための敵対的攻撃(adversarial attacks for good)」と呼ぶ。学習済みモデルへの攻撃として長年研究されてきた摂動や構造化信号は、代わりにデータ所有者、作成者、プラットフォーム、または監査者によって、不正な自動化を妨害したり、その後の説明責任を支援したりするために適用される。5つの研究コミュニティは、ほぼ独立にこの逆転に到達しており、それぞれが視覚資産のライフサイクルの異なる段階を扱っている。すなわち、共有時における不要な認識を防ぐプライバシーフィルタ、不正な学習を防ぐ学習不能な例、悪意のある編集や模倣を防ぐ生成的防御策、自動エージェントに対するアクセス制御のための敵対的CAPTCHA、そして流通後の帰属のための来歴メカニズムである。これらは互いに相容れない成功基準を持つ別々の場で開発されてきたが、これらの手法の多くは、人間の知覚、意味解釈、機械の推論の間に存在し続けるギャップを利用しており、視覚パイプラインがマルチモーダルモデルや自律エージェントへと進化しても、このパラダイムが引き続き関連性を持つことを示唆している。これらの主張を比較可能にするため、我々は5つのファミリーすべてを、転移可能性、適応可能性、展開準備性という共通の軸に沿って評価する。ライフサイクル全体を通して、ほとんどの保護策は依然として静的または適応性の低い攻撃者を主な対象として検証されており、管理されたベンチマークを超えた証拠はほとんどないことが分かる。最後に、堅牢で、構成可能で、展開可能な所有者側保護を実現するためのステージ横断的な対策と未解決問題を整理する。
English
Once visual content enters an AI pipeline, its owner often retains little technical control over how it is used. Legal and regulatory remedies can address misuse, but many technical interventions must be applied earlier, when content is released or accessed. This survey examines the protective paradigm that has grown around this intervention point, which we call adversarial attacks for good. Perturbations and structured signals long studied as attacks on learned models are instead applied by data owners, creators, platforms, or auditors to disrupt unauthorized automation or support later accountability. Five research communities have arrived at this inversion largely independently, each addressing a different stage of a visual asset's lifecycle: privacy filters against unwanted recognition at sharing time, unlearnable examples against unauthorized training, generative safeguards against malicious editing or imitation, adversarial CAPTCHAs for access control against automated agents, and provenance mechanisms for post-circulation attribution. Although developed in separate venues with incompatible success criteria, many of these methods exploit persistent gaps between human perception, semantic interpretation, and machine inference, suggesting that the paradigm remains relevant as visual pipelines evolve toward multimodal models and autonomous agents. To make their claims comparable, we evaluate all five families along shared axes of transferability, adaptability, and deployment readiness. Across the lifecycle, we find that most protections are still validated mainly against static or weakly adaptive adversaries, while evidence beyond controlled benchmarks remains scarce. We close by consolidating cross-stage countermeasures and open problems for robust, composable, and deployable owner-side protection.