OneEmo: 感情知覚・理解・インタラクションのための統合マルチモーダル推論モデル
OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction
August 6, 2026
著者: Jiahao Huang, Zheng Lian, Jingyi Zhang, Zhide Chen, Xiaojiang Peng, Shaonan Wang
cs.AI
要旨
マルチモーダル大規模言語モデル(MLLMs)は、感情知能において顕著な能力を示してきた。しかしながら、既存研究は主にタスク特化型の専門化に焦点を当てており、タスク間の相乗効果を軽視し、潜在的な推論能力を未開拓のままにしていることが多い。このギャップを埋めるために、我々は感情の知覚・理解・相互作用を習得できる統合的感情汎用モデルであるOneEmoを提案する。この目的のため、まず人間参加型ワークフローを通じて専門的な感情知識を明示的な推論軌跡へと蒸留する包括的データセットEmoWorld-130Kを構築する。このコーパスに対する教師ありファインチューニングにより、マルチタスク学習から得られる顕著な相互利益が明らかになる。次に、潜在的な推論能力を完全に引き出すために、統合的なマルチタスク報酬配分を通じて最適化を安定化させる新規の強化学習戦略であるEmo-Chordを提案する。広範な実験により、OneEmoはほとんどのベンチマークにおいて類似サイズのベースラインに対して最先端の性能を達成することを実証する。特筆すべきことに、OneEmoは商用モデルよりもパラメータ数が大幅に少ないにもかかわらず、非常に競争力のある結果をもたらす。本論文は、より信頼性が高く解釈可能な感情コンピューティングへの道を開くものである。コードはhttps://github.com/waHAHJIAHAO/OneEmoで公開している。
English
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld-130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human-in-the-loop workflow. Supervised fine-tuning on this corpus reveals significant mutual benefits derived from multi-task learning. Second, to fully unlock the latent reasoning potential, we propose Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation. Extensive experiments demonstrate that OneEmo achieves state-of-the-art performance against similarly sized baselines across most benchmarks. Notably, despite having significantly fewer parameters than commercial models, OneEmo delivers highly competitive results. This paper paves the way for more reliable and interpretable affective computing. The code is available at https://github.com/waHAHJIAHAO/OneEmo.