Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present <strong>GigaBrain-0.7</strong>, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including 𝜋0.5, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. </p>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"highlights\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#highlights\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tHighlights\n\t</span>\n</h2>\n<ul>\n<li><strong>1. Proposing System-3:</strong> Integrates a world model into the robot’s real-time decision loop, enabling it to <strong>simulate and evaluate before acting.</strong></li>\n<li><strong>2. Ready Out of the Box:</strong> The dual-pyramid framework enables a single pretrained model to perform many tasks, demonstrating the <strong>\"One Model, Many Tasks\"</strong> capability that characterizes foundation models.</li>\n<li><strong>3. Decisive Benchmark Leadership:</strong> GigaBrain-0.7 <strong>leads task success rates</strong> by a wide margin on <strong>Maker H01</strong>, our purpose-built embodied robot, and <strong>ranks first in all four evaluations</strong> on <strong>RoboColiseum</strong>, a platform with strong sim-to-real alignment.</li>\n<li><strong>4. Real-World Tasks in One Continuous Take:</strong> An industry-first demonstration of <strong>over 20 minutes</strong> of long-horizon, complex precision operations spanning <strong>more than 10 tasks</strong>, all in one continuous take.</li>\n</ul>\n<h2 class=\"relative group flex items-baseline\">\n\t<a id=\"resources\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#resources\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tResources\n\t</span>\n</h2>\n<ul>\n<li>🌐 Project: <a href=\"https://gigaai.cc/blog/gigabrain07\" rel=\"nofollow\">https://gigaai.cc/blog/gigabrain07</a></li>\n<li>📄 Paper: <a href=\"https://arxiv.org/abs/2608.15875\" rel=\"nofollow\">https://arxiv.org/abs/2608.15875</a></li>\n<li>💻 Code: <a href=\"https://github.com/open-gigaai/giga-brain-0\" rel=\"nofollow\">https://github.com/open-gigaai/giga-brain-0</a></li>\n<li>🤗 Model: <a href=\"https://huggingface.co/open-gigaai/GigaBrain-0.7-3.5B-Base\">https://huggingface.co/open-gigaai/GigaBrain-0.7-3.5B-Base</a></li>\n<li>🤗 Data: <a href=\"https://huggingface.co/datasets/open-gigaai/GigaBrain-0.7-SampleData\">https://huggingface.co/datasets/open-gigaai/GigaBrain-0.7-SampleData</a></li>\n</ul>\n","updatedAt":"2026-08-26T13:58:07.730Z","author":{"_id":"644012cf3e0374802e174f7c","avatarUrl":"/avatars/0f4a4bd6f96ce193871843e1d01439e8.svg","fullname":"Yang Wang","name":"supermodelteam","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8202372193336487},"editors":["supermodelteam"],"editorAvatarUrls":["/avatars/0f4a4bd6f96ce193871843e1d01439e8.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.15875","authors":[{"_id":"6a8be8d53d26296ea3091a58","name":"GigaBrain Team","hidden":false},{"_id":"6a8be8d53d26296ea3091a59","name":"Angen Ye","hidden":false},{"_id":"6a8be8d53d26296ea3091a5a","name":"Axiang Sun","hidden":false},{"_id":"6a8be8d53d26296ea3091a5b","name":"Can Jin","hidden":false},{"_id":"6a8be8d53d26296ea3091a5c","name":"Chenxi Cheng","hidden":false},{"_id":"6a8be8d53d26296ea3091a5d","name":"Chong Shi","hidden":false},{"_id":"6a8be8d53d26296ea3091a5e","name":"Dengke Shang","hidden":false},{"_id":"6a8be8d53d26296ea3091a5f","name":"Dingqian Zhang","hidden":false},{"_id":"6a8be8d53d26296ea3091a60","name":"Guan Huang","hidden":false},{"_id":"6a8be8d53d26296ea3091a61","name":"Guangqiang Wang","hidden":false},{"_id":"6a8be8d53d26296ea3091a62","user":{"_id":"65dbf86bab2f64915c58b9f4","avatarUrl":"/avatars/f458efccc5b722dbef9c1509ee825edf.svg","isPro":false,"fullname":"Guangqing Ding","user":"abcdgq","type":"user","name":"abcdgq"},"name":"Guangqing Ding","status":"claimed_verified","statusLastChangedAt":"2026-08-26T08:45:04.435Z","hidden":false},{"_id":"6a8be8d53d26296ea3091a63","name":"Guo Li","hidden":false},{"_id":"6a8be8d53d26296ea3091a64","name":"Hangcong Li","hidden":false},{"_id":"6a8be8d53d26296ea3091a65","name":"Hengyu Zhong","hidden":false},{"_id":"6a8be8d53d26296ea3091a66","name":"Hongtao Lu","hidden":false},{"_id":"6a8be8d53d26296ea3091a67","name":"Jianbo Qin","hidden":false},{"_id":"6a8be8d53d26296ea3091a68","name":"Jiming Mao","hidden":false},{"_id":"6a8be8d53d26296ea3091a69","name":"Jing Zhu","hidden":false},{"_id":"6a8be8d53d26296ea3091a6a","name":"Jindi Lv","hidden":false},{"_id":"6a8be8d53d26296ea3091a6b","name":"Jingzhi Cui","hidden":false},{"_id":"6a8be8d53d26296ea3091a6c","name":"Junjie Xie","hidden":false},{"_id":"6a8be8d53d26296ea3091a6d","name":"Junyi Bao","hidden":false},{"_id":"6a8be8d53d26296ea3091a6e","name":"Kai Liu","hidden":false},{"_id":"6a8be8d53d26296ea3091a6f","name":"Lei Yuan","hidden":false},{"_id":"6a8be8d53d26296ea3091a70","name":"Limin Long","hidden":false},{"_id":"6a8be8d53d26296ea3091a71","name":"Lv Feng","hidden":false},{"_id":"6a8be8d53d26296ea3091a72","name":"Mingming Yu","hidden":false},{"_id":"6a8be8d53d26296ea3091a73","name":"Peng Li","hidden":false},{"_id":"6a8be8d53d26296ea3091a74","name":"Pengfei Yi","hidden":false},{"_id":"6a8be8d53d26296ea3091a75","name":"Qi Li","hidden":false},{"_id":"6a8be8d53d26296ea3091a76","name":"Qianli Zhang","hidden":false},{"_id":"6a8be8d53d26296ea3091a77","name":"Qingfang Li","hidden":false},{"_id":"6a8be8d53d26296ea3091a78","name":"Qitang Hu","hidden":false},{"_id":"6a8be8d53d26296ea3091a79","name":"Rui Zhang","hidden":false},{"_id":"6a8be8d53d26296ea3091a7a","name":"Shaoyan Sun","hidden":false},{"_id":"6a8be8d53d26296ea3091a7b","name":"Shibo Sun","hidden":false},{"_id":"6a8be8d53d26296ea3091a7c","name":"Shiying Duan","hidden":false},{"_id":"6a8be8d53d26296ea3091a7d","name":"Tenghui Chen","hidden":false},{"_id":"6a8be8d53d26296ea3091a7e","name":"Tianze Liu","hidden":false},{"_id":"6a8be8d53d26296ea3091a7f","name":"Weijie Ke","hidden":false},{"_id":"6a8be8d53d26296ea3091a80","name":"Wenyao Xue","hidden":false},{"_id":"6a8be8d53d26296ea3091a81","name":"Xiaofeng Wang","hidden":false},{"_id":"6a8be8d53d26296ea3091a82","name":"Xiaoyu Tian","hidden":false},{"_id":"6a8be8d53d26296ea3091a83","name":"Xinyu Liu","hidden":false},{"_id":"6a8be8d53d26296ea3091a84","name":"Xinze Chen","hidden":false},{"_id":"6a8be8d53d26296ea3091a85","user":{"_id":"644012cf3e0374802e174f7c","avatarUrl":"/avatars/0f4a4bd6f96ce193871843e1d01439e8.svg","isPro":false,"fullname":"Yang Wang","user":"supermodelteam","type":"user","name":"supermodelteam"},"name":"Yang Wang","status":"claimed_verified","statusLastChangedAt":"2026-08-26T16:45:04.468Z","hidden":false},{"_id":"6a8be8d53d26296ea3091a86","name":"Yankai Wang","hidden":false},{"_id":"6a8be8d53d26296ea3091a87","name":"Yejun Zeng","hidden":false},{"_id":"6a8be8d53d26296ea3091a88","name":"Yifan Li","hidden":false},{"_id":"6a8be8d53d26296ea3091a89","name":"Yifei Nie","hidden":false},{"_id":"6a8be8d53d26296ea3091a8a","name":"Yilong Li","hidden":false},{"_id":"6a8be8d53d26296ea3091a8b","name":"Yilong Liu","hidden":false},{"_id":"6a8be8d53d26296ea3091a8c","name":"Yongchao Feng","hidden":false},{"_id":"6a8be8d53d26296ea3091a8d","name":"Yumeng Wang","hidden":false},{"_id":"6a8be8d53d26296ea3091a8e","name":"Yun Ye","hidden":false},{"_id":"6a8be8d53d26296ea3091a8f","name":"Zhichao Liu","hidden":false},{"_id":"6a8be8d53d26296ea3091a90","name":"Ziheng He","hidden":false},{"_id":"6a8be8d53d26296ea3091a91","name":"Zonghai Yang","hidden":false},{"_id":"6a8be8d53d26296ea3091a92","name":"Zheng Zhu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/644012cf3e0374802e174f7c/XI8ehgpDazq4TT_NIz0AT.mp4"],"publishedAt":"2026-08-16T00:00:00.000Z","submittedOnDailyAt":"2026-08-26T00:00:00.000Z","title":"GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture","submittedOnDailyBy":{"_id":"644012cf3e0374802e174f7c","avatarUrl":"/avatars/0f4a4bd6f96ce193871843e1d01439e8.svg","isPro":false,"fullname":"Yang Wang","user":"supermodelteam","type":"user","name":"supermodelteam"},"summary":"Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including π_{0.5}, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.","upvotes":92,"discussionId":"6a8be8d53d26296ea3091a93","projectPage":"https://gigaai.cc/blog/gigabrain07","githubRepo":"https://github.com/open-gigaai/giga-brain-0","githubRepoAddedBy":"user","ai_summary":"GigaBrain-0.7 is a vision-language-action model that improves embodied generalization via a three-system architecture, large-scale heterogeneous pretraining, and joint alignment training.","ai_keywords":["vision-language-action models","embodied foundation model","three-system architecture","one-stage alignment training","multi-embodiment action generation","zero-shot capabilities"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2611,"organization":{"_id":"68d6587936e2de9610d9f5f0","name":"open-gigaai","fullname":"GigaAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d6394328e169473e90e4a6/zUK7FKr_8XqrN0aFUgsD-.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6426616ea5ec4a5cbc535634","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6426616ea5ec4a5cbc535634/5IfSFYd9QOxz8K9QmBCst.png","isPro":false,"fullname":"JeffWang","user":"Jeff-Wang","type":"user"},{"_id":"644012cf3e0374802e174f7c","avatarUrl":"/avatars/0f4a4bd6f96ce193871843e1d01439e8.svg","isPro":false,"fullname":"Yang Wang","user":"supermodelteam","type":"user"},{"_id":"6948ee54147afca431ba1d9c","avatarUrl":"/avatars/e7372d2b034737f3f284716ad4db3dcd.svg","isPro":false,"fullname":"li","user":"sayhi886","type":"user"},{"_id":"673beff1d08c9c8e3099ec28","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/f9EuT9uZcHzPJrXWIroZi.png","isPro":false,"fullname":"Pengfei Yi","user":"alfie0","type":"user"},{"_id":"68d63c8e2d41f0f63da038d6","avatarUrl":"/avatars/8979c7380e67229961ff77c5b4d5804b.svg","isPro":false,"fullname":"Guangqiang.Wang","user":"Adannnnnnn","type":"user"},{"_id":"6a4c6ce8d0dfffe1452e28cd","avatarUrl":"/avatars/a3eba21621ae792fc6ba9611fbc46409.svg","isPro":false,"fullname":"xiaoyu","user":"ZUANADAS","type":"user"},{"_id":"67a99bcd3d647533cdf5986f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/FQsZ7x5Pq5rgaV6N9XMf9.png","isPro":false,"fullname":"sunia3345","user":"sunia3345","type":"user"},{"_id":"686fc6d16ea5d5fb0a4a5b2d","avatarUrl":"/avatars/d7a793b03951c2a7682167e17deaf8db.svg","isPro":false,"fullname":"chong","user":"Phil1220","type":"user"},{"_id":"68ad4e9b575612c57f98f8ac","avatarUrl":"/avatars/e0defd4875f261ef0cec6e30dc75cfb4.svg","isPro":false,"fullname":"zz","user":"southrough","type":"user"},{"_id":"654babcef8853732606bab78","avatarUrl":"/avatars/d3e21586583ca684fba08fce2fb38bf1.svg","isPro":false,"fullname":"guo","user":"phoenix88","type":"user"},{"_id":"67c516df4a1fae3ac5a8fc60","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67c516df4a1fae3ac5a8fc60/mSGRBYvxsUfIx5_3U0WKk.jpeg","isPro":false,"fullname":"Michael Sun","user":"MichaelSun1001","type":"user"},{"_id":"65dbf86bab2f64915c58b9f4","avatarUrl":"/avatars/f458efccc5b722dbef9c1509ee825edf.svg","isPro":false,"fullname":"Guangqing Ding","user":"abcdgq","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68d6587936e2de9610d9f5f0","name":"open-gigaai","fullname":"GigaAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d6394328e169473e90e4a6/zUK7FKr_8XqrN0aFUgsD-.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.15875.md","query":{}}">
GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
Abstract
GigaBrain-0.7 is a vision-language-action model that improves embodied generalization via a three-system architecture, large-scale heterogeneous pretraining, and joint alignment training.
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including π_{0.5}, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.
Community
This comment has been hidden (marked as Off-Topic) Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including 𝜋0.5, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios.
Highlights
- 1. Proposing System-3: Integrates a world model into the robot’s real-time decision loop, enabling it to simulate and evaluate before acting.
- 2. Ready Out of the Box: The dual-pyramid framework enables a single pretrained model to perform many tasks, demonstrating the "One Model, Many Tasks" capability that characterizes foundation models.
- 3. Decisive Benchmark Leadership: GigaBrain-0.7 leads task success rates by a wide margin on Maker H01, our purpose-built embodied robot, and ranks first in all four evaluations on RoboColiseum, a platform with strong sim-to-real alignment.
- 4. Real-World Tasks in One Continuous Take: An industry-first demonstration of over 20 minutes of long-horizon, complex precision operations spanning more than 10 tasks, all in one continuous take.
Resources
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.