ChatPaper.aiChatPaper

AffectFlow-DINO: 基于条件整流流的不确定性感知多任务情感估计

AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow

July 14, 2026
作者: Salah Eddine Bekhouche, Abdellah Zakaria Sellam, Fadi Dornaika, Abdenour Hadid
cs.AI

摘要

我们提出了AffectFlow-DINO,一个面向第11届ABAW挑战赛的多任务学习系统。该系统在标准确定性架构基础上扩展了条件整流流头部,用以建模自然场景面部行为中固有的模糊性。模型不直接输出单一情感估计值,而是学习条件生成分布,通过蒙特卡洛采样实现具有不确定性感知的一对多预测。该系统能够从静态人脸图像中同时估计连续效价-唤醒度、分类八种面部表情,并检测十二个动作单元。基于冻结的DINOv3 ViT-S/16骨干网络,大量消融研究表明,整流流解码能持续提升确定性预测的性能,尤其在效价-唤醒度估计方面(CCC-V提升+0.058)。我们进一步证明,后验阈值校准无需重新训练即可有效恢复严重不平衡稀有类别的性能(例如,恐惧类别从3.8%提升至33.1%)。结合骨干网络微调与流重调,最终模型取得了P_MTL=1.177的分数,大幅超越了官方挑战基线P_MTL=0.45。
English
We present AffectFlow-DINO, a multi-task learning system for the 11th ABAW challenge that extends a standard deterministic architecture with a conditional rectified-flow head to model the inherent ambiguity of in-the-wild facial behavior. Instead of predicting a single affect estimate, the model learns a conditional generative distribution, enabling uncertainty-aware one-to-many predictions through Monte Carlo sampling. The system jointly estimates continuous valence-arousal, classifies eight facial expressions, and detects twelve Action Units from static face images. Built on a frozen DINOv3 ViT-S/16 backbone, extensive ablation studies show that rectified-flow decoding consistently improves deterministic prediction, particularly for valence-arousal estimation (CCC-V +0.058). We further show that post-hoc threshold calibration effectively recovers performance on severely imbalanced rare classes (e.g., Fear: 3.8% rightarrow 33.1%) without retraining. Combined with backbone fine-tuning and flow retuning, the final model achieves P_{MTL=1.177}, substantially outperforming the official challenge baseline of P_{MTL}=0.45.