AffectFlow-DINO: 조건부 정류 흐름을 통한 불확실성 인식 다중 작업 감정 추정
AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow
July 14, 2026
저자: Salah Eddine Bekhouche, Abdellah Zakaria Sellam, Fadi Dornaika, Abdenour Hadid
cs.AI
초록
저희는 AffectFlow-DINO를 소개합니다. 이는 제11회 ABAW 챌린지를 위한 다중 작업 학습 시스템으로, 표준 결정론적 아키텍처에 조건부 정규화 흐름 헤드를 확장하여 실제 환경에서의 얼굴 행동이 지닌 본질적 모호성을 모델링합니다. 단일 감정 추정치를 예측하는 대신, 모델은 조건부 생성 분포를 학습하여 몬테카를로 샘플링을 통해 불확실성을 고려한 일대다 예측을 가능하게 합니다. 이 시스템은 정적 얼굴 이미지로부터 연속적인 정서가-각성을 추정하고, 8가지 얼굴 표정을 분류하며, 12가지 행동 단위를 탐지합니다. 고정된 DINOv3 ViT-S/16 백본을 기반으로 한 광범위한 절제 연구는 정규화 흐름 디코딩이 결정론적 예측을 일관되게 개선하며, 특히 정서가-각성 추정(CCC-V +0.058)에서 효과적임을 보여줍니다. 또한, 사후 임계값 보정이 재학습 없이 심각하게 불균형한 희소 클래스(예: 공포: 3.8% → 33.1%)에서 성능을 효과적으로 회복함을 입증합니다. 백본 미세 조정 및 흐름 재조정과 결합하여 최종 모델은 P_MTL=1.177을 달성하며, 공식 챌린지 기준선인 P_MTL=0.45를 크게 능가합니다.
English
We present AffectFlow-DINO, a multi-task learning system for the 11th ABAW challenge that extends a standard deterministic architecture with a conditional rectified-flow head to model the inherent ambiguity of in-the-wild facial behavior. Instead of predicting a single affect estimate, the model learns a conditional generative distribution, enabling uncertainty-aware one-to-many predictions through Monte Carlo sampling. The system jointly estimates continuous valence-arousal, classifies eight facial expressions, and detects twelve Action Units from static face images. Built on a frozen DINOv3 ViT-S/16 backbone, extensive ablation studies show that rectified-flow decoding consistently improves deterministic prediction, particularly for valence-arousal estimation (CCC-V +0.058). We further show that post-hoc threshold calibration effectively recovers performance on severely imbalanced rare classes (e.g., Fear: 3.8% rightarrow 33.1%) without retraining. Combined with backbone fine-tuning and flow retuning, the final model achieves P_{MTL=1.177}, substantially outperforming the official challenge baseline of P_{MTL}=0.45.