ChatPaper.aiChatPaper

運用 CLIP 與 DINO:用於可泛化深偽影像偵測之不確定性感知級聯融合網路

Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection

September 7, 2026
作者: Xuechao Zou, Yi Zhou, Kai Li, Shun Zhang, Yuhui Chen, Congyan Lang, Junliang Xing
cs.AI

摘要

日益逼真的被操控與生成人臉及其易取得性,威脅了數位媒體的可信度。為偵測此類偽造,基於視覺基礎模型的深度偽造偵測器已展現出令人期待的效能,但其通常依賴單一預訓練表徵,且易於過擬合特定訓練分布。為提升對未見偽造的泛化能力,我們提出 UCF-Net,一種不確定性感知級聯融合網路,利用 CLIP 的語言對齊語意先驗與 DINO 的自監督視覺結構先驗。UCF-Net 擷取跨 Transformer 深度的階層式特徵,使用逐層專家聚合以自適應地結合各編碼器的多層級線索,並根據熵推導出的不確定性對所得表徵進行加權融合。我們進一步將公開深度偽造資料集彙整為約 400 萬張影像的統一基準,並建立一個獨立的跨生成器評估集,內含來自八個近期生成器的超過 8,000 張人臉影像。在統一基準上,UCF-Net 於域內與跨域評估中,在受評估方法間取得最佳平均 AUC。在跨生成器集合上,其能以有限的目標域資料有效適應,儘管零樣本遷移仍具挑戰性。
English
The growing realism and accessibility of manipulated and generated faces threaten the trustworthiness of digital media. To detect such forgeries, deepfake detectors based on vision foundation models have shown promising performance, but they typically rely on a single pretrained representation and are prone to overfitting to particular training distributions. To improve generalization to unseen forgeries, we propose UCF-Net, an uncertainty-aware cascaded fusion network that harnesses CLIP's language-aligned semantic priors and DINO's self-supervised visual-structure priors. UCF-Net extracts hierarchical features across Transformer depths, uses layer-wise expert aggregation to adaptively combine each encoder's multi-level cues, and performs weighted fusion of the resulting representations based on entropy-derived uncertainty. We further consolidate public deepfake datasets into a unified benchmark of approximately 4M images and construct a separate cross-generator evaluation set with over 8K face images from eight recent generators. On the unified benchmark, UCF-Net achieves the best mean AUC among the evaluated methods in both in-domain and cross-domain evaluations. On the cross-generator set, it adapts effectively with limited target-domain data, although zero-shot transfer remains challenging.