西红柿、土豆与洋葱:质疑人脸呈现攻击检测中人脸的必要性
Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack Detection
August 20, 2026
作者: Guray Ozgur, Fadi Boutros, Naser Damer
cs.AI
摘要
人脸呈现攻击检测(PAD)传统上被表述为一个面部特定问题,尽管打印、重放和重采集过程引入的许多视觉伪影本质上与面部外观并无关联。在本工作中,我们研究了在下游PAD训练过程中不使用人脸的情况下,能否学习到可迁移的PAD表征。为此,我们提出了TPO,一个受控的、不含人脸的呈现攻击数据集,其中包含几乎随机挑选的西红柿、土豆和洋葱的真实、打印及重放记录,采集协议与常规人脸PAD数据集高度相似。使用基于基础模型的PAD架构,我们证明,在四个标准跨数据集人脸PAD基准上,使用TPO训练的检测器平均AUC达到92.70%,优于使用合成人脸训练,并与在真实人脸数据集上训练的模型保持竞争力。相反,在人脸PAD数据集上训练的模型迁移到TPO时始终高于偶然水平,这表明学习到的表征捕获的是呈现过程的特征,而非物体语义。此外,在固定优化预算下,将TPO纳入常规人脸PAD训练持续提升了跨数据集性能,表明无面部数据提供的是互补信息,而非仅仅是额外的训练样本。最后,表征和频率分析提供了进一步证据,证明可迁移的PAD表征无法由单一的频谱伪影解释,而是编码了跨物体类别共享的更丰富的呈现线索。综合来看,这些结果为以下观点提供了实证依据:可迁移的呈现攻击表征可以独立于面部内容进行学习,为隐私保护和身份无关的PAD开发开辟了新机遇。
English
Face presentation attack detection (PAD) is traditionally formulated as a face-specific problem, although many of the visual artifacts introduced by print, replay, and recapture processes are not inherently tied to facial appearance. In this work, we investigate whether transferable PAD representations can be learned without using faces during downstream PAD training. To this end, we introduce TPO, a controlled face-free presentation attack dataset consisting of bona fide, print, and replay recordings of, almost randomly chosen, tomatoes, potatoes, and onions acquired under protocols that closely mirror conventional face PAD datasets. Using a foundation-model-based PAD architecture, we demonstrate that a detector trained on TPO achieves an average AUC of 92.70% across four standard cross-dataset face PAD benchmarks, outperforming training on synthetic faces and remaining competitive with models trained on real face datasets. Conversely, models trained on face PAD datasets transfer consistently above chance to TPO, suggesting that the learned representations capture characteristics of the presentation process rather than object semantics. Furthermore, incorporating TPO into conventional face PAD training consistently improves cross-dataset performance under fixed optimization budgets, indicating that face-free data provides complementary information rather than simply additional training samples. Finally, representation and frequency analyses provide further evidence that transferable PAD representations cannot be explained by a single spectral artifact but instead encode richer presentation cues shared across object categories. Together, these results provide empirical evidence that transferable presentation attack representations can be learned independently of facial content, opening new opportunities for privacy-preserving and identity-independent PAD development.