Hugging Face Daily Papers · · 4 min read

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

The first agentic RL framework for end-to-end 3D world construction tasks}, providing a unified stack: 2,616 annotated 3D assets, 323 seed 3D worlds, 6,828 user queries, an asset retrieval embedding model, an interactive sandbox environment, a dual-constraint verifier, and a unified agent post-training framework (including both SFT and multimodal RL training)</p>\n","updatedAt":"2026-08-18T02:11:00.809Z","author":{"_id":"65d6a6f7654f85ff0baf161f","avatarUrl":"/avatars/4a46f8b0522fa572c122249a9d6526c4.svg","fullname":"Yansong NING","name":"yasNing","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7988077402114868},"editors":["yasNing"],"editorAvatarUrls":["/avatars/4a46f8b0522fa572c122249a9d6526c4.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.15265","authors":[{"_id":"6a83bee3675db694db8cd46a","name":"Yansong Ning","hidden":false},{"_id":"6a83bee3675db694db8cd46b","name":"Jingwen Ye","hidden":false},{"_id":"6a83bee3675db694db8cd46c","name":"Zhongkai Wu","hidden":false},{"_id":"6a83bee3675db694db8cd46d","name":"Yang Sun","hidden":false},{"_id":"6a83bee3675db694db8cd46e","name":"Yiqin Zhu","hidden":false},{"_id":"6a83bee3675db694db8cd46f","name":"Xingyi Li","hidden":false},{"_id":"6a83bee3675db694db8cd470","name":"Weidong Zhang","hidden":false},{"_id":"6a83bee3675db694db8cd471","name":"Hao Liu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/65d6a6f7654f85ff0baf161f/Vl0S_afut32qbOPXTvnBo.gif"],"publishedAt":"2026-08-15T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?","submittedOnDailyBy":{"_id":"65d6a6f7654f85ff0baf161f","avatarUrl":"/avatars/4a46f8b0522fa572c122249a9d6526c4.svg","isPro":false,"fullname":"Yansong NING","user":"yasNing","type":"user","name":"yasNing"},"summary":"Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user intent, use 3D tools, and reason over textual and visual 3D world information. To this end, we propose VibeWorlding, a unified framework for benchmarking and training vibe worlding agents: a multimodal agent that can autonomously infer user intent, plan scene layout, invoke 3D tools, and reflect on the multimodal feedback in a multi-turn agent-environment interaction process. To achieve this, we first build VWE-BENCH, a benchmark of 2,616 high-quality 3D assets, 323 human-annotated seed 3D worlds, and 6,828 reverse-synthesized multimodal user queries, split into verified queries with ground-truth and unverified queries with carefully designed rubrics. Moreover, we develop VibeWorlding-Gym, a joint multimodal RL post-training framework that integrates (1) a sandbox environment unifying asset retrieval, editing, and image rendering as MCP tools, and (2) a rubric-based verifier that combines physical feasibility and intent fulfillment verification, supporting both fair model evaluation and scalable multimodal RL reward service. Our experiments show that current frontier MLLMs are far from solving the vibe worlding agent task, with even GPT-5.5 and Qwen3.8-Max reaching below 60% success rate, and trace the bottleneck to precise 3D world editing. We further find that RL training can ease this weakness and enable open-source MLLMs to even surpass closed-source frontiers: our VibeWorlder-8B is comparable to frontier MLLMs, while our flagship VibeWorlder-30B-A3B attains the best overall Pass@1 among all evaluated models.","upvotes":41,"discussionId":"6a83bee4675db694db8cd472","projectPage":"https://huggingface.co/collections/usail-hkust/vibeworlder","githubRepo":"https://github.com/usail-hkust/VibeWorlding-Gym","githubRepoAddedBy":"user","ai_summary":"A unified framework benchmarks and trains multimodal agents that infer intent, plan 3D scenes, invoke tools, and reflect on feedback, revealing that reinforcement learning improves open-source models beyond closed-source frontiers.","ai_keywords":["multimodal agents","3D world generation","VibeWorlding","VWE-BENCH","MCP tools","rubric-based verifier","multimodal RL","MLLMs","VibeWorlder"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":24,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65d6a6f7654f85ff0baf161f","avatarUrl":"/avatars/4a46f8b0522fa572c122249a9d6526c4.svg","isPro":false,"fullname":"Yansong NING","user":"yasNing","type":"user"},{"_id":"68d65e99d862fb5b2cad9d3a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/PLJmhy2-dfQAMkbcfDW3A.png","isPro":false,"fullname":"Zherui Yang","user":"yangzhr","type":"user"},{"_id":"66864ab61af627989fddbcbd","avatarUrl":"/avatars/ff532f8bf5af3220f2ae9e778b83fe3a.svg","isPro":false,"fullname":"Wu","user":"Traccc","type":"user"},{"_id":"691dad02a98e73a2c00b7345","avatarUrl":"/avatars/9d8157b5dc40476c2277eb1d50f4e38d.svg","isPro":false,"fullname":"Ziyi Liu","user":"z1ya0","type":"user"},{"_id":"658247c592b5a9664de63882","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658247c592b5a9664de63882/Jn03voLQjDlB3YjpSi-PI.jpeg","isPro":true,"fullname":"fan","user":"EasonFan","type":"user"},{"_id":"65275fe5c3e4d201162df4ed","avatarUrl":"/avatars/4d97aa72864646476930cf7bddb0d0bc.svg","isPro":false,"fullname":"Weilin Lin","user":"Weilin0","type":"user"},{"_id":"64bcc4449f94ea2554edb7c2","avatarUrl":"/avatars/2549e39be03fbab09a680844477191c8.svg","isPro":false,"fullname":"Bowen Liu","user":"liubw","type":"user"},{"_id":"65dd6206e49018d07ba80d59","avatarUrl":"/avatars/16780b6739f83f22230a8b3aa20f6a7f.svg","isPro":false,"fullname":"Wenzhao Jiang","user":"Ciao-Jerry","type":"user"},{"_id":"66e564e8dba1e4fee4572c16","avatarUrl":"/avatars/25bc9694b19f642a4c8c2eab1e1a0de4.svg","isPro":false,"fullname":"Zhixi Chen","user":"EinsamChan","type":"user"},{"_id":"655a02de3fc79984748971da","avatarUrl":"/avatars/e6e70ba38e486ca4695044dda3a16243.svg","isPro":false,"fullname":"Albert Yin","user":"AlbertYin","type":"user"},{"_id":"6746b1e2224b22ef67fbff11","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6746b1e2224b22ef67fbff11/q6sN8lQjtLGCgen_41VxU.jpeg","isPro":false,"fullname":"Zhuoning Guo","user":"Zhuoning","type":"user"},{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"},"query":{}}">
Papers
arxiv:2608.15265

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

Published on Aug 15
· Submitted by
Yansong NING
on Aug 18
#2 Paper of the day
Authors:
,

Abstract

A unified framework benchmarks and trains multimodal agents that infer intent, plan 3D scenes, invoke tools, and reflect on feedback, revealing that reinforcement learning improves open-source models beyond closed-source frontiers.

Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user intent, use 3D tools, and reason over textual and visual 3D world information. To this end, we propose VibeWorlding, a unified framework for benchmarking and training vibe worlding agents: a multimodal agent that can autonomously infer user intent, plan scene layout, invoke 3D tools, and reflect on the multimodal feedback in a multi-turn agent-environment interaction process. To achieve this, we first build VWE-BENCH, a benchmark of 2,616 high-quality 3D assets, 323 human-annotated seed 3D worlds, and 6,828 reverse-synthesized multimodal user queries, split into verified queries with ground-truth and unverified queries with carefully designed rubrics. Moreover, we develop VibeWorlding-Gym, a joint multimodal RL post-training framework that integrates (1) a sandbox environment unifying asset retrieval, editing, and image rendering as MCP tools, and (2) a rubric-based verifier that combines physical feasibility and intent fulfillment verification, supporting both fair model evaluation and scalable multimodal RL reward service. Our experiments show that current frontier MLLMs are far from solving the vibe worlding agent task, with even GPT-5.5 and Qwen3.8-Max reaching below 60% success rate, and trace the bottleneck to precise 3D world editing. We further find that RL training can ease this weakness and enable open-source MLLMs to even surpass closed-source frontiers: our VibeWorlder-8B is comparable to frontier MLLMs, while our flagship VibeWorlder-30B-A3B attains the best overall Pass@1 among all evaluated models.

Community

Paper submitter about 7 hours ago

The first agentic RL framework for end-to-end 3D world construction tasks}, providing a unified stack: 2,616 annotated 3D assets, 323 seed 3D worlds, 6,828 user queries, an asset retrieval embedding model, an interactive sandbox environment, a dual-constraint verifier, and a unified agent post-training framework (including both SFT and multimodal RL training)

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.15265 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers