ChatPaper.aiChatPaper

OneEmo:一种用于情感感知、理解与交互的统一多模态推理模型

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

August 6, 2026
作者: Jiahao Huang, Zheng Lian, Jingyi Zhang, Zhide Chen, Xiaojiang Peng, Shaonan Wang
cs.AI

摘要

多模态大语言模型(MLLMs)在情感智能方面已展现出卓越的能力。然而,现有研究主要集中于任务特定的专门化,常常忽视任务间的协同效应,使潜在推理能力未被充分挖掘。为弥合这一差距,我们提出了OneEmo,一个统一的情感通才模型,能够掌握情绪感知、理解与交互。为此,我们首先构建了EmoWorld-130K,一个综合性数据集,通过人在回路的流程将专门的情感知识提炼为显式推理轨迹。在该语料库上进行监督微调揭示了多任务学习带来的显著互惠效益。其次,为充分释放潜在推理能力,我们提出了Emo-Chord,一种新颖的强化学习策略,通过统一的多任务奖励分配稳定优化过程。大量实验表明,在大多数基准测试中,OneEmo相较于同等规模的基线模型取得了最先进的性能。值得注意的是,尽管参数数量远少于商业模型,OneEmo仍取得了极具竞争力的结果。本文为更可靠、可解释的情感计算铺平了道路。代码可在 https://github.com/waHAHJIAHAO/OneEmo 获取。
English
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld-130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human-in-the-loop workflow. Supervised fine-tuning on this corpus reveals significant mutual benefits derived from multi-task learning. Second, to fully unlock the latent reasoning potential, we propose Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation. Extensive experiments demonstrate that OneEmo achieves state-of-the-art performance against similarly sized baselines across most benchmarks. Notably, despite having significantly fewer parameters than commercial models, OneEmo delivers highly competitive results. This paper paves the way for more reliable and interpretable affective computing. The code is available at https://github.com/waHAHJIAHAO/OneEmo.