Hugging Face Daily Papers · · 5 min read

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Efficient embodied navigation requires more than accurate action selection: an agent must also know when to reason, what to remember, and how to learn from environmental feedback. We introduce TAMP-Nav, where the model selects target points directly in images and projects them into 3D coordinates for execution by a SLAM system. It triggers Chain-of-Thought only at critical decision points, while Anchor-Trajectory Memory preserves important states and compresses routine paths. Two-Level GRPO further combines step-level feedback with trajectory-level outcomes to jointly optimize action selection and selective reasoning. Trained on only 90k trajectories, TAMP-Nav achieves success rates of 66.2% on R2R-CE and 65.7% on RxR-CE while reducing CoT calls by 73.7%. The resulting policy transfers zero-shot to real-world indoor and outdoor environments without real-robot fine-tuning.</p>\n","updatedAt":"2026-08-19T03:38:59.947Z","author":{"_id":"6485bd278d14bcd5cdbb7c8d","avatarUrl":"/avatars/1427cf1a72b5db0cb263ad45885cf925.svg","fullname":"Wenqi Zhang","name":"zwq2018","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":15,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9115067720413208},"editors":["zwq2018"],"editorAvatarUrls":["/avatars/1427cf1a72b5db0cb263ad45885cf925.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.17512","authors":[{"_id":"6a850f2b536bdd3bdd48f720","name":"Hongyan Feng","hidden":false},{"_id":"6a850f2b536bdd3bdd48f721","name":"Sunlai Chen","hidden":false},{"_id":"6a850f2b536bdd3bdd48f722","name":"Xuanyu Liu","hidden":false},{"_id":"6a850f2b536bdd3bdd48f723","name":"Miao Pan","hidden":false},{"_id":"6a850f2b536bdd3bdd48f724","name":"Yangfan Xie","hidden":false},{"_id":"6a850f2b536bdd3bdd48f725","name":"Yuxiang Cui","hidden":false},{"_id":"6a850f2b536bdd3bdd48f726","name":"Zhongxiang Zhou","hidden":false},{"_id":"6a850f2b536bdd3bdd48f727","name":"Rong Xiong","hidden":false},{"_id":"6a850f2b536bdd3bdd48f728","name":"Wenqi Zhang","hidden":false},{"_id":"6a850f2b536bdd3bdd48f729","name":"Jianwei Yin","hidden":false},{"_id":"6a850f2b536bdd3bdd48f72a","name":"Yueting Zhuang","hidden":false},{"_id":"6a850f2b536bdd3bdd48f72b","name":"Xuhong Zhang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6485bd278d14bcd5cdbb7c8d/wZj3rDiw-O9PLeHL-GZZW.mp4"],"publishedAt":"2026-08-18T00:00:00.000Z","submittedOnDailyAt":"2026-08-19T00:00:00.000Z","title":"Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation","submittedOnDailyBy":{"_id":"6485bd278d14bcd5cdbb7c8d","avatarUrl":"/avatars/1427cf1a72b5db0cb263ad45885cf925.svg","isPro":false,"fullname":"Wenqi Zhang","user":"zwq2018","type":"user","name":"zwq2018"},"summary":"Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) that reformulates navigation into 2D visual prompting. Specifically, the VLM merely selects 2D pixels, which are then projected into 3D coordinates for a low-level SLAM controller. This design naturally aligns embodied execution with the VLM's inherent 2D visual capabilities. Second, we propose an integrated Selective Reasoning and Anchor-Trajectory Memory mechanism (Think and Memorize), which dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception. Finally, we design an efficient Two-Level Alignment Paradigm (Align) via Group Relative Policy Optimization (GRPO). By superimposing global outcome rewards with fine-grained process rewards, this dense supervision tightly aligns the agent's cognitive planning with physical environmental feedback, endowing the model with adaptive reasoning capabilities. Experiments demonstrate that TAMP-Nav achieves state-of-the-art performance (e.g., 66.2% SR on R2R-CE) with high runtime and sample efficiency (requiring only 90k training trajectories).","upvotes":29,"discussionId":"6a850f2b536bdd3bdd48f72c","projectPage":"https://zju-omniai.github.io/Embodied-Navigator/","githubRepo":"https://github.com/ZJU-OmniAI/Embodied-Omni","githubRepoAddedBy":"user","ai_summary":"TAMP-Nav improves embodied navigation by aligning vision-language models with 2D visual prompting, selective reasoning with compressed memory, and dense policy optimization.","ai_keywords":["Large Vision-Language Models","Pixel-to-3D Action Formulation","2D visual prompting","SLAM controller","Chain-of-Thought","Anchor-Trajectory Memory","Space-Time Indicators","Group Relative Policy Optimization","global outcome rewards","process rewards"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":208,"organization":{"_id":"696461ab2d94e9a07cdb8efd","name":"OmniAI-ZJU","fullname":"ZJU-OmniAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65f2595830354d0ee043b25a/eEeRdHlGyJ148JQ6BAJ4O.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6485bd278d14bcd5cdbb7c8d","avatarUrl":"/avatars/1427cf1a72b5db0cb263ad45885cf925.svg","isPro":false,"fullname":"Wenqi Zhang","user":"zwq2018","type":"user"},{"_id":"690f009eb9a679e969ece716","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/690f009eb9a679e969ece716/tDU-tAo1JyHKo__e7PLgv.jpeg","isPro":false,"fullname":"xuwang","user":"wx91726","type":"user"},{"_id":"66f10acd6e89e9486f977892","avatarUrl":"/avatars/fe92244b43cde1c2eb759551b8a1de85.svg","isPro":false,"fullname":"Accelerater_C","user":"Accelerater","type":"user"},{"_id":"68ca385a049f422eb68ac0a8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/o0XgN0ebrQ2FAfyAh_6XF.png","isPro":false,"fullname":"kagakouko","user":"kagakouko","type":"user"},{"_id":"6a850d158813aba98a7dca47","avatarUrl":"/avatars/5748c4f7cbaf336eb59250a18816b5b4.svg","isPro":false,"fullname":"HongyanFeng","user":"HongyanFeng","type":"user"},{"_id":"6a8530d6089d0e2b13fd593d","avatarUrl":"/avatars/c57892de50aa424ab61f3d8bcf3ae82c.svg","isPro":false,"fullname":"FHYSmile","user":"FHYSmile","type":"user"},{"_id":"6a8531fb2a19daf99e26a615","avatarUrl":"/avatars/a49cea8a0e1ffed132b4b519eeae98f7.svg","isPro":false,"fullname":"la0820la","user":"la0820la","type":"user"},{"_id":"6a8533402b1dec5ea443c7c0","avatarUrl":"/avatars/42f431ee53d5d2c2a3b125004d2668ba.svg","isPro":false,"fullname":"djfuahva","user":"djfuahva","type":"user"},{"_id":"6a8533b094afb715a9cbec88","avatarUrl":"/avatars/2b435533a834b35144d5d07e93cdd278.svg","isPro":false,"fullname":"vjkash","user":"vjkash","type":"user"},{"_id":"6a8534232a19daf99e26bd5f","avatarUrl":"/avatars/e98d8a21f41285f0e68d6ecd138d86f0.svg","isPro":false,"fullname":"kehqiedv","user":"kehqiedvn","type":"user"},{"_id":"6a8534cf0c5afd86c7d70f2d","avatarUrl":"/avatars/c911e5eb14df287a62a137d968d9576b.svg","isPro":false,"fullname":"mnvakj","user":"mnvakj","type":"user"},{"_id":"6358b570aff68f72ac06113b","avatarUrl":"/avatars/c1df6bab08bc1a940b0af88a3f1e58ad.svg","isPro":false,"fullname":"DawnChase","user":"DawnChase","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"696461ab2d94e9a07cdb8efd","name":"OmniAI-ZJU","fullname":"ZJU-OmniAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65f2595830354d0ee043b25a/eEeRdHlGyJ148JQ6BAJ4O.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.17512.md","query":{}}">
Papers
arxiv:2608.17512

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

Published on Aug 18
· Submitted by
Wenqi Zhang
on Aug 19
#2 Paper of the day
Authors:
,

Abstract

TAMP-Nav improves embodied navigation by aligning vision-language models with 2D visual prompting, selective reasoning with compressed memory, and dense policy optimization.

Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) that reformulates navigation into 2D visual prompting. Specifically, the VLM merely selects 2D pixels, which are then projected into 3D coordinates for a low-level SLAM controller. This design naturally aligns embodied execution with the VLM's inherent 2D visual capabilities. Second, we propose an integrated Selective Reasoning and Anchor-Trajectory Memory mechanism (Think and Memorize), which dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception. Finally, we design an efficient Two-Level Alignment Paradigm (Align) via Group Relative Policy Optimization (GRPO). By superimposing global outcome rewards with fine-grained process rewards, this dense supervision tightly aligns the agent's cognitive planning with physical environmental feedback, endowing the model with adaptive reasoning capabilities. Experiments demonstrate that TAMP-Nav achieves state-of-the-art performance (e.g., 66.2% SR on R2R-CE) with high runtime and sample efficiency (requiring only 90k training trajectories).

Community

Paper submitter about 4 hours ago

Efficient embodied navigation requires more than accurate action selection: an agent must also know when to reason, what to remember, and how to learn from environmental feedback. We introduce TAMP-Nav, where the model selects target points directly in images and projects them into 3D coordinates for execution by a SLAM system. It triggers Chain-of-Thought only at critical decision points, while Anchor-Trajectory Memory preserves important states and compresses routine paths. Two-Level GRPO further combines step-level feedback with trajectory-level outcomes to jointly optimize action selection and selective reasoning. Trained on only 90k trajectories, TAMP-Nav achieves success rates of 66.2% on R2R-CE and 65.7% on RxR-CE while reducing CoT calls by 73.7%. The resulting policy transfers zero-shot to real-world indoor and outdoor environments without real-robot fine-tuning.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.17512
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.17512 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.17512 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.17512 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers