Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) models integrate perception, language, and control, but most represent language as a single global instruction and do not provide an explicit mechanism for assigning and composing arm-specific behaviors. This design limits transfer to collaboration patterns that differ from those observed during training. We present MA-VLA, a unified framework for multi-arm collaboration via atomic action assignment. MA-VLA decomposes cooperative behavior into mid-level atomic prompts and allocates them to individual arms, enabling explicit subgoal specification and compositional reuse across tasks. To reduce reliance on fixed execution roles, we introduce Arm Shuffle, a training-time permutation of the observation, state, and assigned atomic prompts for each arm. This permutation enforces role-agnostic instruction following and supports recomposition into unseen coordination patterns, which we term multi-arm compositional generalization. We also construct a benchmark in which test-time collaboration patterns are absent in training set. Across simulation and real-world evaluations, prior state-of-the-art VLAs largely fail under these unseen collaborations, while MA-VLA consistently succeeds. These results indicate that structured, per-arm atomic action assignment offers a practical route to scalable generalization in multi-arm embodied systems. Code, models, and data are available at <a href=\"https://github.com/zhangzaibin/future-robots\" rel=\"nofollow\">https://github.com/zhangzaibin/future-robots</a></p>\n","updatedAt":"2026-08-27T03:56:14.915Z","author":{"_id":"6575702b15b1ca184b0b2700","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6575702b15b1ca184b0b2700/O9cEodqQmG-gyqMiO_edR.jpeg","fullname":"Zaibin Zhang","name":"MrBean2024","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9045146703720093},"editors":["MrBean2024"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6575702b15b1ca184b0b2700/O9cEodqQmG-gyqMiO_edR.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.25864","authors":[{"_id":"6a8fade92c24e8c5fab329a3","name":"Zaibin Zhang","hidden":false},{"_id":"6a8fade92c24e8c5fab329a4","name":"Junlan Xiao","hidden":false},{"_id":"6a8fade92c24e8c5fab329a5","name":"Zhongbo Zhang","hidden":false},{"_id":"6a8fade92c24e8c5fab329a6","name":"Yifan Wang","hidden":false},{"_id":"6a8fade92c24e8c5fab329a7","name":"Li Kang","hidden":false},{"_id":"6a8fade92c24e8c5fab329a8","name":"Yiran Qin","hidden":false},{"_id":"6a8fade92c24e8c5fab329a9","name":"Changxing Xia","hidden":false},{"_id":"6a8fade92c24e8c5fab329aa","name":"Heng Zhou","hidden":false},{"_id":"6a8fade92c24e8c5fab329ab","name":"Talas Fu","hidden":false},{"_id":"6a8fade92c24e8c5fab329ac","name":"Enshen Zhou","hidden":false},{"_id":"6a8fade92c24e8c5fab329ad","name":"Ruimao Zhang","hidden":false},{"_id":"6a8fade92c24e8c5fab329ae","name":"Zhenfei Yin","hidden":false},{"_id":"6a8fade92c24e8c5fab329af","name":"Huchuan Lu","hidden":false},{"_id":"6a8fade92c24e8c5fab329b0","name":"Lijun Wang","hidden":false}],"publishedAt":"2026-08-26T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization","submittedOnDailyBy":{"_id":"6575702b15b1ca184b0b2700","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6575702b15b1ca184b0b2700/O9cEodqQmG-gyqMiO_edR.jpeg","isPro":false,"fullname":"Zaibin Zhang","user":"MrBean2024","type":"user","name":"MrBean2024"},"summary":"Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) models integrate perception, language, and control, but most represent language as a single global instruction and do not provide an explicit mechanism for assigning and composing arm-specific behaviors. This design limits transfer to collaboration patterns that differ from those observed during training. We present MA-VLA, a unified framework for multi-arm collaboration via atomic action assignment. MA-VLA decomposes cooperative behavior into mid-level atomic prompts and allocates them to individual arms, enabling explicit subgoal specification and compositional reuse across tasks. To reduce reliance on fixed execution roles, we introduce Arm Shuffle, a training-time permutation of the observation, state, and assigned atomic prompts for each arm. This permutation enforces role-agnostic instruction following and supports recomposition into unseen coordination patterns, which we term multi-arm compositional generalization. We also construct a benchmark in which test-time collaboration patterns are absent in training set. Across simulation and real-world evaluations, prior state-of-the-art VLAs largely fail under these unseen collaborations, while MA-VLA consistently succeeds. These results indicate that structured, per-arm atomic action assignment offers a practical route to scalable generalization in multi-arm embodied systems. Code, models, and data are available at https://github.com/zhangzaibin/future-robots","upvotes":2,"discussionId":"6a8fade92c24e8c5fab329b1","projectPage":"https://github.com/zhangzaibin/future-robots","githubRepo":"https://github.com/zhangzaibin/future-robots","githubRepoAddedBy":"user","ai_summary":"MA-VLA enables multi-arm collaboration by assigning atomic actions to individual arms and using training-time permutations to generalize to unseen coordination patterns.","ai_keywords":["vision-language-action models","atomic action assignment","Arm Shuffle","multi-arm compositional generalization","embodied manipulation"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"68b830782b6b404b8318fe8e","name":"dalian-university-of-technology","fullname":"DaLian University of Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68b82cf0116141793335f750/N5laKTgqcFB6x_i8kzdTY.webp"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6575702b15b1ca184b0b2700","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6575702b15b1ca184b0b2700/O9cEodqQmG-gyqMiO_edR.jpeg","isPro":false,"fullname":"Zaibin Zhang","user":"MrBean2024","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68b830782b6b404b8318fe8e","name":"dalian-university-of-technology","fullname":"DaLian University of Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68b82cf0116141793335f750/N5laKTgqcFB6x_i8kzdTY.webp"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.25864.md","query":{}}">
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
Abstract
MA-VLA enables multi-arm collaboration by assigning atomic actions to individual arms and using training-time permutations to generalize to unseen coordination patterns.
Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) models integrate perception, language, and control, but most represent language as a single global instruction and do not provide an explicit mechanism for assigning and composing arm-specific behaviors. This design limits transfer to collaboration patterns that differ from those observed during training. We present MA-VLA, a unified framework for multi-arm collaboration via atomic action assignment. MA-VLA decomposes cooperative behavior into mid-level atomic prompts and allocates them to individual arms, enabling explicit subgoal specification and compositional reuse across tasks. To reduce reliance on fixed execution roles, we introduce Arm Shuffle, a training-time permutation of the observation, state, and assigned atomic prompts for each arm. This permutation enforces role-agnostic instruction following and supports recomposition into unseen coordination patterns, which we term multi-arm compositional generalization. We also construct a benchmark in which test-time collaboration patterns are absent in training set. Across simulation and real-world evaluations, prior state-of-the-art VLAs largely fail under these unseen collaborations, while MA-VLA consistently succeeds. These results indicate that structured, per-arm atomic action assignment offers a practical route to scalable generalization in multi-arm embodied systems. Code, models, and data are available at https://github.com/zhangzaibin/future-robots
Community
Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) models integrate perception, language, and control, but most represent language as a single global instruction and do not provide an explicit mechanism for assigning and composing arm-specific behaviors. This design limits transfer to collaboration patterns that differ from those observed during training. We present MA-VLA, a unified framework for multi-arm collaboration via atomic action assignment. MA-VLA decomposes cooperative behavior into mid-level atomic prompts and allocates them to individual arms, enabling explicit subgoal specification and compositional reuse across tasks. To reduce reliance on fixed execution roles, we introduce Arm Shuffle, a training-time permutation of the observation, state, and assigned atomic prompts for each arm. This permutation enforces role-agnostic instruction following and supports recomposition into unseen coordination patterns, which we term multi-arm compositional generalization. We also construct a benchmark in which test-time collaboration patterns are absent in training set. Across simulation and real-world evaluations, prior state-of-the-art VLAs largely fail under these unseen collaborations, while MA-VLA consistently succeeds. These results indicate that structured, per-arm atomic action assignment offers a practical route to scalable generalization in multi-arm embodied systems. Code, models, and data are available at https://github.com/zhangzaibin/future-robots
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.25864 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.25864 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.25864 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.