<strong>Memory foundation for real-time voice interaction.</strong><br>Project page: <a href=\"https://github.com/xzf-thu/VoiceMem\" rel=\"nofollow\">https://github.com/xzf-thu/VoiceMem</a><br>Code: <a href=\"https://github.com/xzf-thu/VoiceMem\" rel=\"nofollow\">https://github.com/xzf-thu/VoiceMem</a><br>Model: <a href=\"https://huggingface.co/zhifeixie/VoiceMem_MF_Qwen3_6_35B_A3B_Qlora\">https://huggingface.co/zhifeixie/VoiceMem_MF_Qwen3_6_35B_A3B_Qlora</a><br>Dataset: <a href=\"https://huggingface.co/datasets/zhifeixie/VoiceMem-ChatMem400k\">https://huggingface.co/datasets/zhifeixie/VoiceMem-ChatMem400k</a></p>\n","updatedAt":"2026-08-27T04:21:36.220Z","author":{"_id":"678f850bec882f210c1b59f2","avatarUrl":"/avatars/f25bab4afb86de92f75bf9a90e02f59f.svg","fullname":"XIE ZHIFEI","name":"zhifeixie","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":27,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5461278557777405},"editors":["zhifeixie"],"editorAvatarUrls":["/avatars/f25bab4afb86de92f75bf9a90e02f59f.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.26005","authors":[{"_id":"6a8fb8a22c24e8c5fab32a3f","name":"Zhifei Xie","hidden":false},{"_id":"6a8fb8a22c24e8c5fab32a40","name":"Jiaqi Lang","hidden":false},{"_id":"6a8fb8a22c24e8c5fab32a41","name":"Ze An","hidden":false},{"_id":"6a8fb8a22c24e8c5fab32a42","name":"Yifan Zhao","hidden":false},{"_id":"6a8fb8a22c24e8c5fab32a43","name":"Dongchao Yang","hidden":false},{"_id":"6a8fb8a22c24e8c5fab32a44","name":"Kai Li","hidden":false},{"_id":"6a8fb8a22c24e8c5fab32a45","name":"Ziyang Ma","hidden":false},{"_id":"6a8fb8a22c24e8c5fab32a46","name":"Mingbao Lin","hidden":false},{"_id":"6a8fb8a22c24e8c5fab32a47","name":"Chunyan Miao","hidden":false},{"_id":"6a8fb8a22c24e8c5fab32a48","name":"Shuicheng Yan","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/678f850bec882f210c1b59f2/VABXIY3wC7tqlbmduIDw7.gif"],"publishedAt":"2026-08-26T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction","submittedOnDailyBy":{"_id":"678f850bec882f210c1b59f2","avatarUrl":"/avatars/f25bab4afb86de92f75bf9a90e02f59f.svg","isPro":false,"fullname":"XIE ZHIFEI","user":"zhifeixie","type":"user","name":"zhifeixie"},"summary":"Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.","upvotes":77,"discussionId":"6a8fb8a22c24e8c5fab32a49","projectPage":"https://xzf-thu.github.io/VoiceMem/","githubRepo":"https://github.com/xzf-thu/VoiceMem","githubRepoAddedBy":"user","ai_summary":"VoiceMem introduces a dual-brain streaming memory architecture for speech language models that improves retrieval accuracy, emotional personalization, and real-time efficiency.","ai_keywords":["duplex speech language models","VoiceMem","informational left brain","emotional right brain","streaming memory I/O","memory-aware training","affective attribution","dual-node persona modeling","retrieval latency","VAD"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":28,"organization":{"_id":"6371470aafbe42caa5a76208","name":"nanyang-technological-university-singapore","fullname":"Nanyang Technological University Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/637146c5afbe42caa5a75e1b/sZyHSA1AQaAS4nrGan682.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"678f850bec882f210c1b59f2","avatarUrl":"/avatars/f25bab4afb86de92f75bf9a90e02f59f.svg","isPro":false,"fullname":"XIE ZHIFEI","user":"zhifeixie","type":"user"},{"_id":"69afbd376e731c6a5f0b1a63","avatarUrl":"/avatars/41bdc53ff544dfe38293462580750b94.svg","isPro":true,"fullname":"Jiaqi Lang","user":"LangJiaqi77","type":"user"},{"_id":"69e5d357f8405c2e79118e03","avatarUrl":"/avatars/8869f618a7442c42e6dc96686ef03ecf.svg","isPro":false,"fullname":"Emilia","user":"DrEmilia","type":"user"},{"_id":"6916d088f82788e699d7b757","avatarUrl":"/avatars/ab10bdd109e86d7b4ad0afafb48a8d72.svg","isPro":false,"fullname":"fyyyy","user":"zleibston","type":"user"},{"_id":"6387676c23da90491eb9fb16","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1669818175965-noauth.jpeg","isPro":true,"fullname":"Kai Li","user":"JusperLee","type":"user"},{"_id":"6617af2beab5eef6b1e8bb9e","avatarUrl":"/avatars/d939c02027916331d4c44119565f2ca6.svg","isPro":false,"fullname":"XavierJiezou","user":"XavierJiezou","type":"user"},{"_id":"6873b5acee25fbb9bbdba2ee","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/7gmf-pT_pKIcgN5K_ESwF.png","isPro":false,"fullname":"庞锴钰","user":"pangkaiyu","type":"user"},{"_id":"6a8fc28fefa404d0730aa77a","avatarUrl":"/avatars/7010aa08addb3c4cd04f323e959ee9cb.svg","isPro":false,"fullname":"Strain042","user":"Strain042","type":"user"},{"_id":"6a80a5ae0e81b9f49c576b16","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a80a5ae0e81b9f49c576b16/gLVRZ0F9CbRRJ790xyoTL.jpeg","isPro":false,"fullname":"Justin Rahman","user":"justinrahman","type":"user"},{"_id":"6a87de51714cb7eb8b868391","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a87de51714cb7eb8b868391/yTpUNznBQrevHZKGyZpTl.jpeg","isPro":false,"fullname":"Bruna P. Alves","user":"brunaalves0827","type":"user"},{"_id":"6a7e064cffe2488635f97bc6","avatarUrl":"/avatars/0d25bc7030dc8d1b71122540ff4c03bd.svg","isPro":false,"fullname":"Noah Davis","user":"ivankuznetsovdud","type":"user"},{"_id":"6a8b7094553594c49ff3a8b8","avatarUrl":"/avatars/fb1f65b38ab28898a9082fbf9139dc4b.svg","isPro":false,"fullname":"प्रिया कुमार","user":"priyanair","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"6371470aafbe42caa5a76208","name":"nanyang-technological-university-singapore","fullname":"Nanyang Technological University Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/637146c5afbe42caa5a75e1b/sZyHSA1AQaAS4nrGan682.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.26005.md","query":{}}">
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Abstract
VoiceMem introduces a dual-brain streaming memory architecture for speech language models that improves retrieval accuracy, emotional personalization, and real-time efficiency.
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.26005 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.26005 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.26005 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.