VoiceMem: リアルタイムインタラクションのためのストリーミング型デュアルブレインメモリ
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
August 26, 2026
著者: Zhifei Xie, Jiaqi Lang, Ze An, Yifan Zhao, Dongchao Yang, Kai Li, Ziyang Ma, Mingbao Lin, Chunyan Miao, Shuicheng Yan
cs.AI
要旨
デュプレックス音声言語モデル(SLM)などの対話システムには、その中核として、ストリーミング対応かつ正確で共感的なメモリシステムが依然として欠けている。我々はVoiceMemを提案する。これは、並列的な情報処理を担う「左脳」、感情を司る「右脳」、そしてストリーミング式のメモリI/O機構を備えたシンプルなメモリアーキテクチャである。さらに、メモリ認識型SLMトレーニング、長期評価、交換可能なメモリバックエンドによる分離型デプロイメントのための完全なパイプラインも構築した。
実験および実環境への導入により、以下の3つの利点が実証された。i) 正確性:Top-5検索の条件下で、左脳はMem0などの従来システムのTop-200検索を約30ポイント上回る性能を達成した。ii) 感情的・個人的側面:右脳は、短期・長期の感情帰属とデュアルノードのペルソナモデリングにより、3つのペルソナベンチマークすべてで最先端の性能を達成し、従来の最良システムに対し総合スコアを4.29ポイント向上させた。iii) リアルタイム性と低コスト:VoiceMemは検索を134msで完了し、これは標準的なVADレイテンシに十分収まる。そのため追加の対話遅延を生じさせず、高い精度と低コストを維持する。
これらの結果は、VoiceMemがリアルタイムでパーソナライズされ、感情を認識する音声対話のための実用的なメモリ基盤を提供することを示している。
English
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.