ChatPaper.aiChatPaper

토마토, 감자, 양파: 얼굴 프레젠테이션 공격 탐지에서 얼굴의 필요성에 대한 의문 제기

Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack Detection

August 20, 2026
저자: Guray Ozgur, Fadi Boutros, Naser Damer
cs.AI

초록

얼굴 제시 공격 탐지(PAD)는 전통적으로 얼굴 특정 문제로 정식화되지만, 인쇄, 재생, 재촬영 과정에서 도입되는 많은 시각적 인공물들은 본질적으로 얼굴 외관과 관련되지 않는다. 본 연구에서는 다운스트림 PAD 훈련 중 얼굴을 사용하지 않고도 전이 가능한 PAD 표현을 학습할 수 있는지 조사한다. 이를 위해 우리는 기존 얼굴 PAD 데이터셋을 밀접하게 반영하는 프로토콜 하에 획득된, 거의 무작위로 선택된 토마토, 감자, 양파의 실제(bona fide), 인쇄, 재생 녹화물로 구성된 통제된 얼굴 미포함 제시 공격 데이터셋 TPO를 소개한다. 기반 모델(foundation model) 기반 PAD 아키텍처를 사용하여, TPO로 훈련된 탐지기가 네 가지 표준 데이터셋 간 얼굴 PAD 벤치마크에서 평균 AUC 92.70%를 달성하며, 합성 얼굴 훈련을 능가하고 실제 얼굴 데이터셋으로 훈련된 모델과 경쟁력을 유지함을 입증한다. 반대로, 얼굴 PAD 데이터셋으로 훈련된 모델은 TPO에 대해 일관되게 우연 수준 이상으로 전이되며, 이는 학습된 표현이 객체 의미론(object semantics)보다는 제시 과정의 특성을 포착함을 시사한다. 더욱이, TPO를 기존 얼굴 PAD 훈련에 통합하면 고정된 최적화 예산 하에서 데이터셋 간 성능이 일관되게 향상되며, 이는 얼굴이 없는 데이터가 단순한 추가 훈련 샘플이 아니라 상보적 정보를 제공함을 나타낸다. 마지막으로, 표현 및 주파수 분석은 전이 가능한 PAD 표현이 단일 스펙트럼 인공물로 설명될 수 없으며, 대신 객체 범주 간에 공유되는 더 풍부한 제시 단서를 인코딩한다는 추가 증거를 제공한다. 종합하면, 이러한 결과는 전이 가능한 제시 공격 표현이 얼굴 콘텐츠와 무관하게 학습될 수 있다는 경험적 증거를 제공하며, 프라이버시를 보호하고 신원에 독립적인 PAD 개발의 새로운 기회를 열어준다.
English
Face presentation attack detection (PAD) is traditionally formulated as a face-specific problem, although many of the visual artifacts introduced by print, replay, and recapture processes are not inherently tied to facial appearance. In this work, we investigate whether transferable PAD representations can be learned without using faces during downstream PAD training. To this end, we introduce TPO, a controlled face-free presentation attack dataset consisting of bona fide, print, and replay recordings of, almost randomly chosen, tomatoes, potatoes, and onions acquired under protocols that closely mirror conventional face PAD datasets. Using a foundation-model-based PAD architecture, we demonstrate that a detector trained on TPO achieves an average AUC of 92.70% across four standard cross-dataset face PAD benchmarks, outperforming training on synthetic faces and remaining competitive with models trained on real face datasets. Conversely, models trained on face PAD datasets transfer consistently above chance to TPO, suggesting that the learned representations capture characteristics of the presentation process rather than object semantics. Furthermore, incorporating TPO into conventional face PAD training consistently improves cross-dataset performance under fixed optimization budgets, indicating that face-free data provides complementary information rather than simply additional training samples. Finally, representation and frequency analyses provide further evidence that transferable PAD representations cannot be explained by a single spectral artifact but instead encode richer presentation cues shared across object categories. Together, these results provide empirical evidence that transferable presentation attack representations can be learned independently of facial content, opening new opportunities for privacy-preserving and identity-independent PAD development.