Efficient embodied navigation requires more than accurate action selection: an agent must also know when to reason, what to remember, and how to learn from environmental feedback. We introduce TAMP-Nav, where the model selects target points directly in images and projects them into 3D coordinates for execution by a SLAM system. It triggers Chain-of-Thought only at critical decision points, while Anchor-Trajectory Memory preserves important states and compresses routine paths. Two-Level GRPO further combines step-level feedback with trajectory-level outcomes to jointly optimize action selection and selective reasoning. Trained on only 90k trajectories, TAMP-Nav achieves success rates of 66.2% on R2R-CE and 65.7% on RxR-CE while reducing CoT calls by 73.7%. The resulting policy transfers zero-shot to real-world indoor and outdoor environments without real-robot fine-tuning.</p>\n","updatedAt":"2026-08-19T03:38:59.947Z","author":{"_id":"6485bd278d14bcd5cdbb7c8d","avatarUrl":"/avatars/1427cf1a72b5db0cb263ad45885cf925.svg","fullname":"Wenqi Zhang","name":"zwq2018","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":15,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9115067720413208},"editors":["zwq2018"],"editorAvatarUrls":["/avatars/1427cf1a72b5db0cb263ad45885cf925.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.17512","authors":[{"_id":"6a850f2b536bdd3bdd48f720","name":"Hongyan Feng","hidden":false},{"_id":"6a850f2b536bdd3bdd48f721","name":"Sunlai Chen","hidden":false},{"_id":"6a850f2b536bdd3bdd48f722","name":"Xuanyu Liu","hidden":false},{"_id":"6a850f2b536bdd3bdd48f723","name":"Miao Pan","hidden":false},{"_id":"6a850f2b536bdd3bdd48f724","name":"Yangfan Xie","hidden":false},{"_id":"6a850f2b536bdd3bdd48f725","name":"Yuxiang Cui","hidden":false},{"_id":"6a850f2b536bdd3bdd48f726","name":"Zhongxiang Zhou","hidden":false},{"_id":"6a850f2b536bdd3bdd48f727","name":"Rong Xiong","hidden":false},{"_id":"6a850f2b536bdd3bdd48f728","name":"Wenqi Zhang","hidden":false},{"_id":"6a850f2b536bdd3bdd48f729","name":"Jianwei Yin","hidden":false},{"_id":"6a850f2b536bdd3bdd48f72a","name":"Yueting Zhuang","hidden":false},{"_id":"6a850f2b536bdd3bdd48f72b","name":"Xuhong Zhang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6485bd278d14bcd5cdbb7c8d/wZj3rDiw-O9PLeHL-GZZW.mp4"],"publishedAt":"2026-08-18T00:00:00.000Z","submittedOnDailyAt":"2026-08-19T00:00:00.000Z","title":"Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation","submittedOnDailyBy":{"_id":"6485bd278d14bcd5cdbb7c8d","avatarUrl":"/avatars/1427cf1a72b5db0cb263ad45885cf925.svg","isPro":false,"fullname":"Wenqi Zhang","user":"zwq2018","type":"user","name":"zwq2018"},"summary":"Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) that reformulates navigation into 2D visual prompting. Specifically, the VLM merely selects 2D pixels, which are then projected into 3D coordinates for a low-level SLAM controller. This design naturally aligns embodied execution with the VLM's inherent 2D visual capabilities. Second, we propose an integrated Selective Reasoning and Anchor-Trajectory Memory mechanism (Think and Memorize), which dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception. Finally, we design an efficient Two-Level Alignment Paradigm (Align) via Group Relative Policy Optimization (GRPO). By superimposing global outcome rewards with fine-grained process rewards, this dense supervision tightly aligns the agent's cognitive planning with physical environmental feedback, endowing the model with adaptive reasoning capabilities. Experiments demonstrate that TAMP-Nav achieves state-of-the-art performance (e.g., 66.2% SR on R2R-CE) with high runtime and sample efficiency (requiring only 90k training trajectories).","upvotes":29,"discussionId":"6a850f2b536bdd3bdd48f72c","projectPage":"https://zju-omniai.github.io/Embodied-Navigator/","githubRepo":"https://github.com/ZJU-OmniAI/Embodied-Omni","githubRepoAddedBy":"user","ai_summary":"TAMP-Nav improves embodied navigation by aligning vision-language models with 2D visual prompting, selective reasoning with compressed memory, and dense policy optimization.","ai_keywords":["Large Vision-Language Models","Pixel-to-3D Action Formulation","2D visual prompting","SLAM controller","Chain-of-Thought","Anchor-Trajectory Memory","Space-Time Indicators","Group Relative Policy Optimization","global outcome rewards","process rewards"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":208,"organization":{"_id":"696461ab2d94e9a07cdb8efd","name":"OmniAI-ZJU","fullname":"ZJU-OmniAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65f2595830354d0ee043b25a/eEeRdHlGyJ148JQ6BAJ4O.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6485bd278d14bcd5cdbb7c8d","avatarUrl":"/avatars/1427cf1a72b5db0cb263ad45885cf925.svg","isPro":false,"fullname":"Wenqi Zhang","user":"zwq2018","type":"user"},{"_id":"690f009eb9a679e969ece716","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/690f009eb9a679e969ece716/tDU-tAo1JyHKo__e7PLgv.jpeg","isPro":false,"fullname":"xuwang","user":"wx91726","type":"user"},{"_id":"66f10acd6e89e9486f977892","avatarUrl":"/avatars/fe92244b43cde1c2eb759551b8a1de85.svg","isPro":false,"fullname":"Accelerater_C","user":"Accelerater","type":"user"},{"_id":"68ca385a049f422eb68ac0a8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/o0XgN0ebrQ2FAfyAh_6XF.png","isPro":false,"fullname":"kagakouko","user":"kagakouko","type":"user"},{"_id":"6a850d158813aba98a7dca47","avatarUrl":"/avatars/5748c4f7cbaf336eb59250a18816b5b4.svg","isPro":false,"fullname":"HongyanFeng","user":"HongyanFeng","type":"user"},{"_id":"6a8530d6089d0e2b13fd593d","avatarUrl":"/avatars/c57892de50aa424ab61f3d8bcf3ae82c.svg","isPro":false,"fullname":"FHYSmile","user":"FHYSmile","type":"user"},{"_id":"6a8531fb2a19daf99e26a615","avatarUrl":"/avatars/a49cea8a0e1ffed132b4b519eeae98f7.svg","isPro":false,"fullname":"la0820la","user":"la0820la","type":"user"},{"_id":"6a8533402b1dec5ea443c7c0","avatarUrl":"/avatars/42f431ee53d5d2c2a3b125004d2668ba.svg","isPro":false,"fullname":"djfuahva","user":"djfuahva","type":"user"},{"_id":"6a8533b094afb715a9cbec88","avatarUrl":"/avatars/2b435533a834b35144d5d07e93cdd278.svg","isPro":false,"fullname":"vjkash","user":"vjkash","type":"user"},{"_id":"6a8534232a19daf99e26bd5f","avatarUrl":"/avatars/e98d8a21f41285f0e68d6ecd138d86f0.svg","isPro":false,"fullname":"kehqiedv","user":"kehqiedvn","type":"user"},{"_id":"6a8534cf0c5afd86c7d70f2d","avatarUrl":"/avatars/c911e5eb14df287a62a137d968d9576b.svg","isPro":false,"fullname":"mnvakj","user":"mnvakj","type":"user"},{"_id":"6358b570aff68f72ac06113b","avatarUrl":"/avatars/c1df6bab08bc1a940b0af88a3f1e58ad.svg","isPro":false,"fullname":"DawnChase","user":"DawnChase","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"696461ab2d94e9a07cdb8efd","name":"OmniAI-ZJU","fullname":"ZJU-OmniAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65f2595830354d0ee043b25a/eEeRdHlGyJ148JQ6BAJ4O.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.17512.md","query":{}}">
Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
Abstract
TAMP-Nav improves embodied navigation by aligning vision-language models with 2D visual prompting, selective reasoning with compressed memory, and dense policy optimization.
Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) that reformulates navigation into 2D visual prompting. Specifically, the VLM merely selects 2D pixels, which are then projected into 3D coordinates for a low-level SLAM controller. This design naturally aligns embodied execution with the VLM's inherent 2D visual capabilities. Second, we propose an integrated Selective Reasoning and Anchor-Trajectory Memory mechanism (Think and Memorize), which dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception. Finally, we design an efficient Two-Level Alignment Paradigm (Align) via Group Relative Policy Optimization (GRPO). By superimposing global outcome rewards with fine-grained process rewards, this dense supervision tightly aligns the agent's cognitive planning with physical environmental feedback, endowing the model with adaptive reasoning capabilities. Experiments demonstrate that TAMP-Nav achieves state-of-the-art performance (e.g., 66.2% SR on R2R-CE) with high runtime and sample efficiency (requiring only 90k training trajectories).
Community
Efficient embodied navigation requires more than accurate action selection: an agent must also know when to reason, what to remember, and how to learn from environmental feedback. We introduce TAMP-Nav, where the model selects target points directly in images and projects them into 3D coordinates for execution by a SLAM system. It triggers Chain-of-Thought only at critical decision points, while Anchor-Trajectory Memory preserves important states and compresses routine paths. Two-Level GRPO further combines step-level feedback with trajectory-level outcomes to jointly optimize action selection and selective reasoning. Trained on only 90k trajectories, TAMP-Nav achieves success rates of 66.2% on R2R-CE and 65.7% on RxR-CE while reducing CoT calls by 73.7%. The resulting policy transfers zero-shot to real-world indoor and outdoor environments without real-robot fine-tuning.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.17512 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.17512 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.17512 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.