CLIP과 DINO 활용: 일반화 가능한 딥페이크 이미지 탐지를 위한 불확실성 인지 캐스케이드 융합 네트워크
Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection
September 7, 2026
저자: Xuechao Zou, Yi Zhou, Kai Li, Shun Zhang, Yuhui Chen, Congyan Lang, Junliang Xing
cs.AI
초록
조작되거나 생성된 얼굴의 사실성과 접근성이 높아짐에 따라 디지털 미디어의 신뢰성이 위협받고 있다. 이러한 위조를 탐지하기 위해 비전 파운데이션 모델 기반 딥페이크 탐지기들이 유망한 성능을 보여 왔지만, 이들은 일반적으로 단일 사전학습 표현에 의존하며 특정 학습 분포에 과적합되기 쉽다. 보지 못한 위조에 대한 일반화를 향상하기 위해, 우리는 CLIP의 언어 정렬 의미 프라이어와 DINO의 자기지도 시각 구조 프라이어를 활용하는 불확실성 인식 캐스케이드 융합 네트워크인 UCF-Net을 제안한다. UCF-Net은 Transformer 깊이 전반에 걸쳐 계층적 특징을 추출하고, 레이어별 전문가 집계를 사용해 각 인코더의 다층 단서를 적응적으로 결합하며, 엔트로피에서 도출된 불확실성에 기반해 결과 표현들의 가중 융합을 수행한다. 또한 우리는 공개 딥페이크 데이터셋들을 약 400만 장 이미지의 통합 벤치마크로 통합하고, 8개의 최신 생성기에서 얻은 8천 장 이상의 얼굴 이미지로 구성된 별도의 교차 생성기 평가 세트를 구축한다. 통합 벤치마크에서 UCF-Net은 평가된 방법들 가운데 도메인 내 및 교차 도메인 평가 모두에서 최고 평균 AUC를 달성한다. 교차 생성기 세트에서는 제한된 타깃 도메인 데이터로 효과적으로 적응하지만, 제로샷 전이는 여전히 어려운 과제로 남아 있다.
English
The growing realism and accessibility of manipulated and generated faces threaten the trustworthiness of digital media. To detect such forgeries, deepfake detectors based on vision foundation models have shown promising performance, but they typically rely on a single pretrained representation and are prone to overfitting to particular training distributions. To improve generalization to unseen forgeries, we propose UCF-Net, an uncertainty-aware cascaded fusion network that harnesses CLIP's language-aligned semantic priors and DINO's self-supervised visual-structure priors. UCF-Net extracts hierarchical features across Transformer depths, uses layer-wise expert aggregation to adaptively combine each encoder's multi-level cues, and performs weighted fusion of the resulting representations based on entropy-derived uncertainty. We further consolidate public deepfake datasets into a unified benchmark of approximately 4M images and construct a separate cross-generator evaluation set with over 8K face images from eight recent generators. On the unified benchmark, UCF-Net achieves the best mean AUC among the evaluated methods in both in-domain and cross-domain evaluations. On the cross-generator set, it adapts effectively with limited target-domain data, although zero-shot transfer remains challenging.