ChatPaper.aiChatPaper

AffectFlow-DINO:基於條件校正流的不確定感知多任務情感估計

AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow

July 14, 2026
作者: Salah Eddine Bekhouche, Abdellah Zakaria Sellam, Fadi Dornaika, Abdenour Hadid
cs.AI

摘要

我們提出 AffectFlow-DINO,這是一個用於第 11 屆 ABAW 挑戰賽的多任務學習系統,透過引入條件整流流頭部來擴展標準的確定性架構,以模擬自然場景下臉部行為的內在模糊性。該模型並非預測單一情感估計值,而是學習條件生成分佈,透過蒙地卡羅取樣實現具有不確定性意識的一對多預測。此系統從靜態臉部影像中聯合估計連續的效價-喚醒度、分類八種臉部表情,並檢測十二個動作單元。基於凍結的 DINOv3 ViT-S/16 骨幹網路,廣泛的消融研究表明,整流流解碼持續改進確定性預測,特別是在效價-喚醒度估計方面(CCC-V 提升 0.058)。我們進一步證明,事後閾值校準能夠在不重新訓練的情況下,有效恢復嚴重不平衡的稀有類別(例如:恐懼從 3.8% 提升至 33.1%)的性能。結合骨幹網路微調與流重調優,最終模型達到 P_MTL=1.177,顯著優於官方挑戰基準線 P_MTL=0.45。
English
We present AffectFlow-DINO, a multi-task learning system for the 11th ABAW challenge that extends a standard deterministic architecture with a conditional rectified-flow head to model the inherent ambiguity of in-the-wild facial behavior. Instead of predicting a single affect estimate, the model learns a conditional generative distribution, enabling uncertainty-aware one-to-many predictions through Monte Carlo sampling. The system jointly estimates continuous valence-arousal, classifies eight facial expressions, and detects twelve Action Units from static face images. Built on a frozen DINOv3 ViT-S/16 backbone, extensive ablation studies show that rectified-flow decoding consistently improves deterministic prediction, particularly for valence-arousal estimation (CCC-V +0.058). We further show that post-hoc threshold calibration effectively recovers performance on severely imbalanced rare classes (e.g., Fear: 3.8% rightarrow 33.1%) without retraining. Combined with backbone fine-tuning and flow retuning, the final model achieves P_{MTL=1.177}, substantially outperforming the official challenge baseline of P_{MTL}=0.45.