Memory is NOT always what you need, as it may impair rather than enhance model capabilities.</p>\n","updatedAt":"2026-08-21T02:13:03.119Z","author":{"_id":"64bf898d979949d2e2585c9a","avatarUrl":"/avatars/da77c856ec997e2b812c06272a01c8b2.svg","fullname":"mengruwang","name":"mengru","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9735316038131714},"editors":["mengru"],"editorAvatarUrls":["/avatars/da77c856ec997e2b812c06272a01c8b2.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.20202","authors":[{"_id":"6a87b2c389e517cbfd75dc01","name":"Mengru Wang","hidden":false},{"_id":"6a87b2c389e517cbfd75dc02","name":"Haozhe Luo","hidden":false},{"_id":"6a87b2c389e517cbfd75dc03","name":"Zhenqian Xu","hidden":false},{"_id":"6a87b2c389e517cbfd75dc04","name":"Zhixiang Cui","hidden":false},{"_id":"6a87b2c389e517cbfd75dc05","name":"Haoming Xu","hidden":false},{"_id":"6a87b2c389e517cbfd75dc06","name":"Qu Yang","hidden":false},{"_id":"6a87b2c389e517cbfd75dc07","name":"Jizhan Fang","hidden":false},{"_id":"6a87b2c389e517cbfd75dc08","name":"Junfeng Fang","hidden":false},{"_id":"6a87b2c389e517cbfd75dc09","name":"Ningyu Zhang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64bf898d979949d2e2585c9a/umUIVtLbmIStJUUpNnA7y.png"],"publishedAt":"2026-08-20T00:00:00.000Z","submittedOnDailyAt":"2026-08-21T00:00:00.000Z","title":"MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use","submittedOnDailyBy":{"_id":"64bf898d979949d2e2585c9a","avatarUrl":"/avatars/da77c856ec997e2b812c06272a01c8b2.svg","isPro":false,"fullname":"mengruwang","user":"mengru","type":"user","name":"mengru"},"summary":"Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memories reshape model reasoning and affect performance on the current task. We identify memory-induced cognitive traps: even faithfully recorded and semantically relevant memories can distort model reasoning or beliefs and degrade current task performance. To systematically evaluate these failure modes, we introduce MemTrapBench, which covers two forms of cognitive traps: Reasoning Fixation and Belief Distortion. Experiments across two model families and five representative memory frameworks show that MemTrapBench is challenging: all evaluated memory strategies underperform the no-memory setting, with even the strongest methods suffering drops of more than 10%. To mitigate these cognitive traps, we propose AdaptiveMem, a simple yet effective inference-time method that instructs LLMs to avoid memory traps. AdaptiveMem mitigates cognitive traps on MemTrapBench while preserving or improving performance on standard memory benchmarks across diverse memory frameworks.","upvotes":24,"discussionId":"6a87b2c389e517cbfd75dc0a","githubRepo":"https://github.com/zjunlp/MemTrapBench","githubRepoAddedBy":"user","ai_summary":"Retrieved memories can induce reasoning errors and belief distortions in large language models, and an inference-time strategy helps avoid these cognitive traps while maintaining benchmark performance.","ai_keywords":["memory-induced cognitive traps","Reasoning Fixation","Belief Distortion","MemTrapBench","AdaptiveMem","inference-time method","memory frameworks"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"620a6fcd8d5e5dfed284bc91","name":"zjunlp","fullname":"ZJUNLP","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1644851027419-620a61cba53066560e226d30.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"672c198760bdd070539fd7ed","avatarUrl":"/avatars/1064f0f5929c505589ee77f3e36df7e9.svg","isPro":false,"fullname":"Bohao Wang","user":"Baymax0110","type":"user"},{"_id":"64bf898d979949d2e2585c9a","avatarUrl":"/avatars/da77c856ec997e2b812c06272a01c8b2.svg","isPro":false,"fullname":"mengruwang","user":"mengru","type":"user"},{"_id":"6846bc1873f604b8d273e04f","avatarUrl":"/avatars/0110b442f893dfc74e2051dfa2e9aee6.svg","isPro":false,"fullname":"Haozhe Luo","user":"Carloooooo","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6190ab805ca89a28e9f66873","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6190ab805ca89a28e9f66873/5sU31QRyrKjL9OA64sMUk.jpeg","isPro":false,"fullname":"Xin Xu","user":"XinXuNLPer","type":"user"},{"_id":"609b320efe087f3d04cf047b","avatarUrl":"/avatars/629869f54e5b7774025fc225783555a3.svg","isPro":false,"fullname":"lilei","user":"flow3rdown","type":"user"},{"_id":"620b3bbb0668e435407c8d0a","avatarUrl":"/avatars/e0fccbb2577d76088e09f054c35cffbc.svg","isPro":true,"fullname":"Ningyu Zhang","user":"Ningyu","type":"user"},{"_id":"6441f1d2603214724ec0c1c2","avatarUrl":"/avatars/d3c4b759e6a5635e37ff715fae52e5ba.svg","isPro":false,"fullname":"Shumin Deng","user":"231sm","type":"user"},{"_id":"6a87bcf3eda6dc52f19e0486","avatarUrl":"/avatars/da4f5477283b84565eda8c0a45a5eccb.svg","isPro":false,"fullname":"wu zesen","user":"zesenwu23","type":"user"},{"_id":"63a942dd2e05ca32e35335df","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63a942dd2e05ca32e35335df/kuKfBLEXfWnvnoUUmoXW6.jpeg","isPro":false,"fullname":"haoming xu","user":"haomingx","type":"user"},{"_id":"652bdbb77c5365f2d1228dfb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/652bdbb77c5365f2d1228dfb/ImPwcK1dMr23MtJVI9C9I.jpeg","isPro":false,"fullname":"ZhongYi","user":"Blurblur02","type":"user"},{"_id":"6776ae0c91b4c75dac91249c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6776ae0c91b4c75dac91249c/uJk3ZnRrzjPCcBNjmrWLI.png","isPro":false,"fullname":"Oran Feng","user":"xiachongfeng","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"620a6fcd8d5e5dfed284bc91","name":"zjunlp","fullname":"ZJUNLP","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1644851027419-620a61cba53066560e226d30.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.20202.md","query":{}}">
MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use
Abstract
Retrieved memories can induce reasoning errors and belief distortions in large language models, and an inference-time strategy helps avoid these cognitive traps while maintaining benchmark performance.
Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memories reshape model reasoning and affect performance on the current task. We identify memory-induced cognitive traps: even faithfully recorded and semantically relevant memories can distort model reasoning or beliefs and degrade current task performance. To systematically evaluate these failure modes, we introduce MemTrapBench, which covers two forms of cognitive traps: Reasoning Fixation and Belief Distortion. Experiments across two model families and five representative memory frameworks show that MemTrapBench is challenging: all evaluated memory strategies underperform the no-memory setting, with even the strongest methods suffering drops of more than 10%. To mitigate these cognitive traps, we propose AdaptiveMem, a simple yet effective inference-time method that instructs LLMs to avoid memory traps. AdaptiveMem mitigates cognitive traps on MemTrapBench while preserving or improving performance on standard memory benchmarks across diverse memory frameworks.
Community
Memory is NOT always what you need, as it may impair rather than enhance model capabilities.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.20202 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.20202 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.20202 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.