ChatPaper.aiChatPaper

AffectFlow-DINO: 条件付き整流フローによる不確実性を考慮したマルチタスク感情推定

AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow

July 14, 2026
著者: Salah Eddine Bekhouche, Abdellah Zakaria Sellam, Fadi Dornaika, Abdenour Hadid
cs.AI

要旨

我々はAffectFlow-DINOを提案する。これは第11回ABAWチャレンジ向けのマルチタスク学習システムであり、標準的な決定論的アーキテクチャに条件付き整流流ヘッドを拡張することで、実環境における顔行動の本質的な曖昧性をモデル化する。単一の感情推定値を予測する代わりに、条件付き生成分布を学習し、モンテカルロサンプリングによる不確実性を考慮したOne-to-Many予測を可能にする。本システムは、静的な顔画像から連続的なバレンス・アラウザルを推定し、8種類の表情を分類し、12種類のアクションユニットを検出する。凍結したDINOv3 ViT-S/16バックボーンを基盤とし、広範なアブレーション研究により、整流流デコードが決定論的予測を一貫して改善し、特にバレンス・アラウザル推定(CCC-V +0.058)で顕著であることを示す。さらに、事後的な閾値調整により、再学習なしで深刻な不均衡な稀なクラス(例:恐怖:3.8%→33.1%)の性能を効果的に回復できることを示す。バックボーンの微調整とフローの再調整を組み合わせた最終モデルは、P_{MTL}=1.177を達成し、公式チャレンジベースラインのP_{MTL}=0.45を大幅に上回る。
English
We present AffectFlow-DINO, a multi-task learning system for the 11th ABAW challenge that extends a standard deterministic architecture with a conditional rectified-flow head to model the inherent ambiguity of in-the-wild facial behavior. Instead of predicting a single affect estimate, the model learns a conditional generative distribution, enabling uncertainty-aware one-to-many predictions through Monte Carlo sampling. The system jointly estimates continuous valence-arousal, classifies eight facial expressions, and detects twelve Action Units from static face images. Built on a frozen DINOv3 ViT-S/16 backbone, extensive ablation studies show that rectified-flow decoding consistently improves deterministic prediction, particularly for valence-arousal estimation (CCC-V +0.058). We further show that post-hoc threshold calibration effectively recovers performance on severely imbalanced rare classes (e.g., Fear: 3.8% rightarrow 33.1%) without retraining. Combined with backbone fine-tuning and flow retuning, the final model achieves P_{MTL=1.177}, substantially outperforming the official challenge baseline of P_{MTL}=0.45.