ChatPaper.aiChatPaper

OneEmo:一個用於情緒感知、理解與互動的統一多模態推理模型

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

August 6, 2026
作者: Jiahao Huang, Zheng Lian, Jingyi Zhang, Zhide Chen, Xiaojiang Peng, Shaonan Wang
cs.AI

摘要

多模態大型語言模型(MLLMs)在情感智能方面已展現出卓越的能力。然而,當前研究主要集中於任務特定的專門化,往往忽略了任務間的協同效應,使得潛在的推理潛力未被充分發掘。為填補這一空白,我們提出 OneEmo,一個統一的情感通用模型,能夠掌握情感感知、理解與互動。為此,我們首先構建了 EmoWorld-130K,一個綜合性資料集,透過人在迴路工作流程將專門的情感知識提煉為明確的推理軌跡。基於該語料庫的監督式微調顯示,多任務學習能帶來顯著的相互效益。其次,為充分釋放潛在推理潛力,我們提出 Emo-Chord,一種新穎的強化學習策略,透過統一的多任務獎勵分配來穩定最佳化過程。大量實驗表明,OneEmo 在大多數基準測試中均達到與同規模基線模型相當的最佳性能。值得注意的是,儘管 OneEmo 的參數量遠少於商業模型,其結果仍極具競爭力。本文為更可靠且可解釋的情感計算開闢了道路。程式碼已公開於 https://github.com/waHAHJIAHAO/OneEmo。
English
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld-130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human-in-the-loop workflow. Supervised fine-tuning on this corpus reveals significant mutual benefits derived from multi-task learning. Second, to fully unlock the latent reasoning potential, we propose Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation. Extensive experiments demonstrate that OneEmo achieves state-of-the-art performance against similarly sized baselines across most benchmarks. Notably, despite having significantly fewer parameters than commercial models, OneEmo delivers highly competitive results. This paper paves the way for more reliable and interpretable affective computing. The code is available at https://github.com/waHAHJIAHAO/OneEmo.