VoiceMem: 실시간 상호작용을 위한 스트리밍 듀얼 브레인 메모리

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

August 26, 2026
저자: Zhifei Xie, Jiaqi Lang, Ze An, Yifan Zhao, Dongchao Yang, Kai Li, Ziyang Ma, Mingbao Lin, Chunyan Miao, Shuicheng Yan
cs.AI

초록

듀플렉스 음성 언어 모델(SLM)과 같은 대화형 시스템은 여전히 그 영혼에 해당하는 스트리밍 방식의 정확하고 공감적인 메모리 시스템을 갖추지 못하고 있다. 우리는 병렬적 정보 처리 좌뇌, 감정 처리 우뇌, 그리고 스트리밍 메모리 입출력 메커니즘을 갖춘 간단한 메모리 아키텍처인 VoiceMem을 제안한다. 또한 메모리 인지 SLM 훈련, 장기 평가, 교체 가능한 메모리 백엔드를 통한 분리형 배포를 위한 완전한 파이프라인을 구축한다. 실험과 실제 배포를 통해 세 가지 장점을 확인했다: i) 정확성: 상위 5개 검색에서 좌뇌는 Mem0과 같은 기존 시스템의 상위 200개 검색 성능보다 거의 30포인트 높은 성능을 보인다; ii) 감정 및 개인화: 단기·장기 정서적 귀인과 이중 노드 페르소나 모델링을 갖춘 우뇌는 세 가지 페르소나 벤치마크에서 최첨단 성능을 달성하며, 이전 최고 시스템 대비 종합 점수가 4.29포인트 향상된다; iii) 실시간 및 저비용: VoiceMem은 134ms 내에 검색을 완료하여 표준 VAD 지연 시간에 충분히 부합하며, 추가 대화 지연 없이 높은 정확도와 낮은 비용을 유지한다. 이러한 결과는 VoiceMem이 실시간 개인화 및 감정 인지 음성 상호작용을 위한 실용적인 메모리 기반을 제공함을 보여준다.
English
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.
PDF1501August 28, 2026