トマト、ジャガイモ、タマネギ:顔提示攻撃検出における顔の必要性に疑問を投げかける
Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack Detection
August 20, 2026
著者: Guray Ozgur, Fadi Boutros, Naser Damer
cs.AI
要旨
顔提示攻撃検出(PAD)は、伝統的に顔固有の問題として定式化されているが、印刷、再生、再撮影のプロセスによって生じる視覚的アーティファクトの多くは、本質的に顔の外観に結びついてはいない。本研究では、下流のPADトレーニング中に顔を使用せずに、転移可能なPAD表現を学習できるかどうかを調査する。この目的のために、我々はTPOを紹介する。TPOは、従来の顔PADデータセットと密接に対応するプロトコルに基づいて取得された、ほぼランダムに選択されたトマト、ジャガイモ、タマネギの正規録画、印刷録画、再生録画からなる、制御された顔を含まない提示攻撃データセットである。ファウンデーションモデルベースのPADアーキテクチャを用いて、TPOでトレーニングされた検出器が、4つの標準的なクロスデータセット顔PADベンチマーク全体で平均AUC 92.70%を達成し、合成顔でのトレーニングを上回り、実顔データセットでトレーニングされたモデルと競合する性能を維持することを実証する。逆に、顔PADデータセットでトレーニングされたモデルは、TPOに対して一貫して偶然レベルを超える転移を示し、学習された表現がオブジェクトの意味論ではなく提示プロセスの特性を捉えていることを示唆している。さらに、従来の顔PADトレーニングにTPOを組み込むと、固定された最適化予算の下でクロスデータセット性能が一貫して向上し、顔を含まないデータが単なる追加のトレーニングサンプルではなく補完的な情報を提供することを示している。最後に、表現分析と周波数分析は、転移可能なPAD表現が単一のスペクトルアーティファクトによって説明されるのではなく、オブジェクトカテゴリ間で共有されるより豊かな提示手がかりを符号化するというさらなる証拠を提供する。これらの結果は総合的に、転移可能な提示攻撃表現が顔の内容から独立して学習され得るという実証的証拠を提供し、プライバシー保護およびアイデンティティ非依存のPAD開発の新たな機会を開くものである。
English
Face presentation attack detection (PAD) is traditionally formulated as a face-specific problem, although many of the visual artifacts introduced by print, replay, and recapture processes are not inherently tied to facial appearance. In this work, we investigate whether transferable PAD representations can be learned without using faces during downstream PAD training. To this end, we introduce TPO, a controlled face-free presentation attack dataset consisting of bona fide, print, and replay recordings of, almost randomly chosen, tomatoes, potatoes, and onions acquired under protocols that closely mirror conventional face PAD datasets. Using a foundation-model-based PAD architecture, we demonstrate that a detector trained on TPO achieves an average AUC of 92.70% across four standard cross-dataset face PAD benchmarks, outperforming training on synthetic faces and remaining competitive with models trained on real face datasets. Conversely, models trained on face PAD datasets transfer consistently above chance to TPO, suggesting that the learned representations capture characteristics of the presentation process rather than object semantics. Furthermore, incorporating TPO into conventional face PAD training consistently improves cross-dataset performance under fixed optimization budgets, indicating that face-free data provides complementary information rather than simply additional training samples. Finally, representation and frequency analyses provide further evidence that transferable PAD representations cannot be explained by a single spectral artifact but instead encode richer presentation cues shared across object categories. Together, these results provide empirical evidence that transferable presentation attack representations can be learned independently of facial content, opening new opportunities for privacy-preserving and identity-independent PAD development.