Spatial Memory Agent introduces a practical, parameter-update-free approach to improving frozen VLMs on spatial reasoning. By distilling verified experience into transferable memory and calibrating retrieval reliability, SMA delivers consistent gains across multiple models and benchmarks without relying on external spatial tools at inference time.</p>\n","updatedAt":"2026-08-14T03:02:50.462Z","author":{"_id":"632179745fc60c44fd91fc33","avatarUrl":"/avatars/37d4fefbcc19f091dccffefec9706de2.svg","fullname":"zhumuzhi","name":"Z-MU-Z","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8492933511734009},"editors":["Z-MU-Z"],"editorAvatarUrls":["/avatars/37d4fefbcc19f091dccffefec9706de2.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.12743","authors":[{"_id":"6a7e719942823931a1f1754c","name":"Haokai Zhang","hidden":false},{"_id":"6a7e719942823931a1f1754d","name":"Yuhang Ding","hidden":false},{"_id":"6a7e719942823931a1f1754e","name":"Yunshu Zhou","hidden":false},{"_id":"6a7e719942823931a1f1754f","name":"Xinze Du","hidden":false},{"_id":"6a7e719942823931a1f17550","user":{"_id":"65e87ca9c3ef3307f6ab0139","avatarUrl":"/avatars/1a8b75c5a4b0d09e8f1980ef4b29e1a1.svg","isPro":false,"fullname":"shengtao zhang","user":"zz2403","type":"user","name":"zz2403"},"name":"Shengtao Zhang","status":"claimed_verified","statusLastChangedAt":"2026-08-14T08:45:04.603Z","hidden":false},{"_id":"6a7e719942823931a1f17551","name":"Zhiyue Zhao","hidden":false},{"_id":"6a7e719942823931a1f17552","name":"Yuling Xi","hidden":false},{"_id":"6a7e719942823931a1f17553","name":"Hao Chen","hidden":false}],"publishedAt":"2026-08-13T00:00:00.000Z","submittedOnDailyAt":"2026-08-14T00:00:00.000Z","title":"Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence","submittedOnDailyBy":{"_id":"632179745fc60c44fd91fc33","avatarUrl":"/avatars/37d4fefbcc19f091dccffefec9706de2.svg","isPro":false,"fullname":"zhumuzhi","user":"Z-MU-Z","type":"user","name":"Z-MU-Z"},"summary":"Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fine-tuning and reinforcement learning. Another line adopts an agentic paradigm in which the model calls external spatial tools, such as depth estimation and 3D reconstruction tools, to gather intermediate spatial evidence. We study a complementary and underexplored route: Can a frozen VLM agent improve its spatial reasoning through parameter-update-free self-evolution, without depending on external expert spatial tools at inference time? We present Spatial Memory Agent (SMA), an experience-grounded runtime framework that converts verified spatial experience into reusable transferable lessons. In a verifiable spatial environment, SMA queries the frozen VLM, obtains a predicted answer and reward, and uses verifier-guided reflection to distill compact transferable lessons from spatial experience. SMA further assigns each lesson a Transfer Reliability Score (TRS), which is initialized uniformly and calibrated from later retrieval outcomes as visit evidence of future transfer reliability. During read-only deployment, SMA retrieves lessons by semantic filter and similarity-TRS combined ranking, allowing the retrieved memory to guide frozen model inference. Across five representative spatial benchmarks and four base VLMs, SMA achieves the highest macro average in every base-model block and the best accuracy among the evaluated methods in most of the 20 evaluations, establishing a practical parameter-update-free path for spatial self-evolution across the evaluated frozen model scales and environments.","upvotes":25,"discussionId":"6a7e719942823931a1f17554","projectPage":"https://aim-uofa.github.io/SMA/","ai_summary":"A frozen vision-language model improves spatial reasoning by self-evolving through verified experience, reflection, and reusable memory retrieval without parameter updates or external tools.","ai_keywords":["parameter-update-free self-evolution","verifier-guided reflection","Transfer Reliability Score","experience-grounded runtime framework","spatial memory agent","read-only deployment"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"61bac2af530e5c78d7b99667","name":"zju","fullname":"Zhejiang University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5e1058e9fcf41d740b69966d/7G1xjlxwCdMEmKcxNR0n5.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6512a55c4151fb1fa722c4f7","avatarUrl":"/avatars/c7cebc47d424c3c81c4c7e8a0f55c713.svg","isPro":false,"fullname":"AI explorer","user":"ToexploreAI","type":"user"},{"_id":"6a0bec00f03495eb57f03041","avatarUrl":"/avatars/50687a2238bed8e2ac2810ef6ece7262.svg","isPro":false,"fullname":"zhouyunshu","user":"dewlumina","type":"user"},{"_id":"698612c46c9345fe291cc8ed","avatarUrl":"/avatars/d05fa66f89ba130b1d2c69fe303a303e.svg","isPro":false,"fullname":"D","user":"baiyeD","type":"user"},{"_id":"681f4a09b82e61bb502fb030","avatarUrl":"/avatars/0da3a829cb2abaec649d58d80814ee27.svg","isPro":false,"fullname":"CodyQ","user":"CodyQ623","type":"user"},{"_id":"6a7e76188b63f3edc9672bc9","avatarUrl":"/avatars/c03496c7eb8213df3ad0f92221572726.svg","isPro":false,"fullname":"Patrick June","user":"Patrick3924","type":"user"},{"_id":"67151d9576da0cd1a863ec53","avatarUrl":"/avatars/9aec66df33b1b87ec7aa3ab77f1c22f1.svg","isPro":false,"fullname":"李翱希","user":"stormycity","type":"user"},{"_id":"65e87ca9c3ef3307f6ab0139","avatarUrl":"/avatars/1a8b75c5a4b0d09e8f1980ef4b29e1a1.svg","isPro":false,"fullname":"shengtao zhang","user":"zz2403","type":"user"},{"_id":"6422c943e2029ade06edf69b","avatarUrl":"/avatars/2a0c19841f32f41f244017ec2ba0ad38.svg","isPro":false,"fullname":"wangjiaqian","user":"jiaqian","type":"user"},{"_id":"6a7e781803a49ba96beb9d3b","avatarUrl":"/avatars/b7813ddfaee7b931665eaf445bc1c006.svg","isPro":false,"fullname":"galaxy","user":"creategalaxysupernova","type":"user"},{"_id":"67bd558d1d5babb0426811a4","avatarUrl":"/avatars/951e63c8b73e59bdaa66a65073848419.svg","isPro":false,"fullname":"x","user":"yulingxi","type":"user"},{"_id":"632179745fc60c44fd91fc33","avatarUrl":"/avatars/37d4fefbcc19f091dccffefec9706de2.svg","isPro":false,"fullname":"zhumuzhi","user":"Z-MU-Z","type":"user"},{"_id":"67496c0baaee270ea7020477","avatarUrl":"/avatars/5c9ed0887eab6526bbbdc03af9ea5feb.svg","isPro":false,"fullname":"hxk","user":"xingyang1","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61bac2af530e5c78d7b99667","name":"zju","fullname":"Zhejiang University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5e1058e9fcf41d740b69966d/7G1xjlxwCdMEmKcxNR0n5.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.12743.md","query":{}}">
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
Abstract
A frozen vision-language model improves spatial reasoning by self-evolving through verified experience, reflection, and reusable memory retrieval without parameter updates or external tools.
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fine-tuning and reinforcement learning. Another line adopts an agentic paradigm in which the model calls external spatial tools, such as depth estimation and 3D reconstruction tools, to gather intermediate spatial evidence. We study a complementary and underexplored route: Can a frozen VLM agent improve its spatial reasoning through parameter-update-free self-evolution, without depending on external expert spatial tools at inference time? We present Spatial Memory Agent (SMA), an experience-grounded runtime framework that converts verified spatial experience into reusable transferable lessons. In a verifiable spatial environment, SMA queries the frozen VLM, obtains a predicted answer and reward, and uses verifier-guided reflection to distill compact transferable lessons from spatial experience. SMA further assigns each lesson a Transfer Reliability Score (TRS), which is initialized uniformly and calibrated from later retrieval outcomes as visit evidence of future transfer reliability. During read-only deployment, SMA retrieves lessons by semantic filter and similarity-TRS combined ranking, allowing the retrieved memory to guide frozen model inference. Across five representative spatial benchmarks and four base VLMs, SMA achieves the highest macro average in every base-model block and the best accuracy among the evaluated methods in most of the 20 evaluations, establishing a practical parameter-update-free path for spatial self-evolution across the evaluated frozen model scales and environments.
Community
Spatial Memory Agent introduces a practical, parameter-update-free approach to improving frozen VLMs on spatial reasoning. By distilling verified experience into transferable memory and calibrating retrieval reliability, SMA delivers consistent gains across multiple models and benchmarks without relying on external spatial tools at inference time.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.12743 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.12743 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.12743 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.