Hugging Face Daily Papers · · 3 min read

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Thanks for InfRec.</p>\n","updatedAt":"2026-08-28T04:37:56.757Z","author":{"_id":"648157807a741c7f33f5ebf5","avatarUrl":"/avatars/57a620971ce3e1ec883dc0772a5fb0b1.svg","fullname":"Pengfei Zhou","name":"IntJudge","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":3,"identifiedLanguage":{"language":"en","probability":0.929682195186615},"editors":["IntJudge"],"editorAvatarUrls":["/avatars/57a620971ce3e1ec883dc0772a5fb0b1.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.25518","authors":[{"_id":"6a910fafa64059bab69c36cf","name":"Pengfei Zhou","hidden":false},{"_id":"6a910fafa64059bab69c36d0","name":"Hexin Wang","hidden":false},{"_id":"6a910fafa64059bab69c36d1","name":"Zhengfeiyang Zhang","hidden":false},{"_id":"6a910fafa64059bab69c36d2","name":"Yixing Ma","hidden":false},{"_id":"6a910fafa64059bab69c36d3","name":"Zhenglin Wan","hidden":false},{"_id":"6a910fafa64059bab69c36d4","name":"Kaipeng Zhang","hidden":false},{"_id":"6a910fafa64059bab69c36d5","name":"Wangbo Zhao","hidden":false},{"_id":"6a910fafa64059bab69c36d6","name":"Yang You","hidden":false}],"publishedAt":"2026-08-26T00:00:00.000Z","submittedOnDailyAt":"2026-08-28T00:00:00.000Z","title":"Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models","submittedOnDailyBy":{"_id":"648157807a741c7f33f5ebf5","avatarUrl":"/avatars/57a620971ce3e1ec883dc0772a5fb0b1.svg","isPro":false,"fullname":"Pengfei Zhou","user":"IntJudge","type":"user","name":"IntJudge"},"summary":"A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers grounded reward signals. The success of code agents illustrates why this matters. As code is executable, compilers and runtimes can provide high-quality rewards for Reinforcement Learning (RL) post-training of LLMs. By contrast, spatial generation still relies largely on fuzzy proxies such as CLIP scores. These signals are fuzzy and biased, making them hard to support RL post-training. Compared with these, game development provides a missing reward environment for spatial world models. A scene encoded by a game engine is an executable world specification: the engine can efficiently check collision, physics, navigability and bounded playability, while the developer provides the global verification signal by judging whether the scene should be accepted. Game development also provides real-world long-horizon trajectory data for RL post-training. We therefore propose Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process.","upvotes":1,"discussionId":"6a910fafa64059bab69c36d7","ai_summary":"Game engines provide executable verification and long-horizon trajectories for reinforcement learning post-training of spatial world models, motivating a human-engine verification paradigm.","ai_keywords":["world models","reinforcement learning","RL post-training","CLIP scores","game engine","executable world specification","collision","physics","navigability","bounded playability","RLHEV","Reinforcement Learning with Human-Engine Verification"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"6508ab2b349930913196378b","name":"NationalUniversityofSingapore","fullname":"National University of Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/630ca0817dacb93b33506ce7/ZYUmpSMsa5Whihw3me2Bw.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"648157807a741c7f33f5ebf5","avatarUrl":"/avatars/57a620971ce3e1ec883dc0772a5fb0b1.svg","isPro":false,"fullname":"Pengfei Zhou","user":"IntJudge","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6508ab2b349930913196378b","name":"NationalUniversityofSingapore","fullname":"National University of Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/630ca0817dacb93b33506ce7/ZYUmpSMsa5Whihw3me2Bw.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.25518.md","query":{}}">
Papers
arxiv:2608.25518

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

Published on Aug 26
· Submitted by
Pengfei Zhou
on Aug 28
Authors:
,

Abstract

Game engines provide executable verification and long-horizon trajectories for reinforcement learning post-training of spatial world models, motivating a human-engine verification paradigm.

A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers grounded reward signals. The success of code agents illustrates why this matters. As code is executable, compilers and runtimes can provide high-quality rewards for Reinforcement Learning (RL) post-training of LLMs. By contrast, spatial generation still relies largely on fuzzy proxies such as CLIP scores. These signals are fuzzy and biased, making them hard to support RL post-training. Compared with these, game development provides a missing reward environment for spatial world models. A scene encoded by a game engine is an executable world specification: the engine can efficiently check collision, physics, navigability and bounded playability, while the developer provides the global verification signal by judging whether the scene should be accepted. Game development also provides real-world long-horizon trajectory data for RL post-training. We therefore propose Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process.

Community

Thanks for InfRec.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.25518
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.25518 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.25518 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.25518 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers