VoiceMem:用於即時互動的串流雙腦記憶

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

August 26, 2026
作者: Zhifei Xie, Jiaqi Lang, Ze An, Yifan Zhao, Dongchao Yang, Kai Li, Ziyang Ma, Mingbao Lin, Chunyan Miao, Shuicheng Yan
cs.AI

摘要

對話系統(例如雙工語音語言模型,SLM)仍缺乏一個流式、準確且具備同理心的記憶系統作為其靈魂。我們提出VoiceMem,這是一個簡單的記憶架構,具備並行的資訊型左腦、情感型右腦,以及流式記憶輸入/輸出機制。我們進一步建構了一套完整的流程,涵蓋記憶感知的SLM訓練、長時程評估,以及採用可替換記憶後端的解耦部署。實驗與真實世界部署展現了三項優勢:i) 準確性:在top-5檢索下,左腦的表現比Mem0等經典系統在top-200時高出近30個百分點;ii) 情感與個人化:右腦結合短程與長程情感歸因及雙節點人格建模,在三個人物性格基準測試中達到最先進的效能,並比先前最佳系統的綜合分數高出4.29分;iii) 即時且低成本:VoiceMem在134毫秒內完成檢索,遠低於標準VAD延遲,在不增加額外對話延遲的同時,維持高準確度與低成本。這些結果顯示VoiceMem為即時、個人化且情感感知的語音互動提供了實用的記憶基礎。
English
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.
PDF1501August 28, 2026