ChatPaper.aiChatPaper

番茄、馬鈴薯與洋蔥:質疑人臉呈現攻擊檢測中人臉之必要性

Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack Detection

August 20, 2026
作者: Guray Ozgur, Fadi Boutros, Naser Damer
cs.AI

摘要

人臉呈現攻擊檢測(PAD)傳統上被視為一個針對人臉的問題,儘管打印、重放和重新捕獲過程引入的許多視覺偽影並非本質上與面部外觀相關。在本工作中,我們研究是否可以在下游PAD訓練中不使用人臉的情況下學習可遷移的PAD表示。為此,我們提出了TPO,這是一個受控的無人臉呈現攻擊數據集,包含對幾乎隨機選擇的番茄、馬鈴薯和洋蔥的真實、打印和重放錄製,其採集協議緊密模仿傳統的人臉PAD數據集。基於基礎模型的PAD架構,我們證明在TPO上訓練的檢測器在四個標準跨數據集人臉PAD基準上平均AUC達到92.70%,優於在合成人臉上的訓練,並且與在真實人臉數據集上訓練的模型保持競爭力。相反,在人臉PAD數據集上訓練的模型遷移到TPO時持續表現優於隨機水平,表明所學習的表示捕捉的是呈現過程的特徵而非物體語義。此外,將TPO納入傳統的人臉PAD訓練中,在固定優化預算下持續提升跨數據集性能,表明無人臉數據提供的是互補信息,而不僅僅是額外的訓練樣本。最後,表示分析和頻率分析提供了進一步的證據,表明可遷移的PAD表示不能由單一的頻譜偽影來解釋,而是編碼了跨物體類別共享的更豐富的呈現線索。綜合來看,這些結果提供了經驗證據,表明可遷移的呈現攻擊表示可以獨立於面部內容進行學習,為保護隱私和與身份無關的PAD開發開闢了新的機會。
English
Face presentation attack detection (PAD) is traditionally formulated as a face-specific problem, although many of the visual artifacts introduced by print, replay, and recapture processes are not inherently tied to facial appearance. In this work, we investigate whether transferable PAD representations can be learned without using faces during downstream PAD training. To this end, we introduce TPO, a controlled face-free presentation attack dataset consisting of bona fide, print, and replay recordings of, almost randomly chosen, tomatoes, potatoes, and onions acquired under protocols that closely mirror conventional face PAD datasets. Using a foundation-model-based PAD architecture, we demonstrate that a detector trained on TPO achieves an average AUC of 92.70% across four standard cross-dataset face PAD benchmarks, outperforming training on synthetic faces and remaining competitive with models trained on real face datasets. Conversely, models trained on face PAD datasets transfer consistently above chance to TPO, suggesting that the learned representations capture characteristics of the presentation process rather than object semantics. Furthermore, incorporating TPO into conventional face PAD training consistently improves cross-dataset performance under fixed optimization budgets, indicating that face-free data provides complementary information rather than simply additional training samples. Finally, representation and frequency analyses provide further evidence that transferable PAD representations cannot be explained by a single spectral artifact but instead encode richer presentation cues shared across object categories. Together, these results provide empirical evidence that transferable presentation attack representations can be learned independently of facial content, opening new opportunities for privacy-preserving and identity-independent PAD development.