VoiceMem:面向实时交互的流式双脑记忆

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

August 26, 2026
作者: Zhifei Xie, Jiaqi Lang, Ze An, Yifan Zhao, Dongchao Yang, Kai Li, Ziyang Ma, Mingbao Lin, Chunyan Miao, Shuicheng Yan
cs.AI

摘要

对话系统(例如双工语音语言模型,SLMs)仍缺少一个流式、准确且富有同理心的记忆系统作为其灵魂。我们提出 VoiceMem,一个简洁的记忆架构,包含并行的信息型左脑、情感型右脑以及流式记忆输入/输出机制。我们进一步构建了一条完整的流水线,涵盖记忆感知的 SLM 训练、长时程评估以及支持可替换记忆后端的解耦部署。实验与真实部署表明其三大优势:i) 准确性:在前 5 检索条件下,左脑相比 Mem0 等经典系统在前 200 检索时领先近 30 个百分点;ii) 情感与个性化:右脑通过短时程与长时程情感归因以及双节点人格建模,在三个人格基准上达到最先进水平,并在综合评分上较此前最优系统提升 4.29 分;iii) 实时与低成本:VoiceMem 在 134 ms 内完成检索,完全处于标准 VAD 延迟范围内,在不引入额外对话延迟的同时保持了高准确率与低成本。这些结果表明,VoiceMem 为实时、个性化且具备情感感知能力的语音交互提供了实用的记忆基础。
English
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.
PDF1501August 28, 2026