We investigate the probabilistic failures in video generation models and present a reliable benchmark to measure it.</p>\n","updatedAt":"2026-08-28T02:41:59.742Z","author":{"_id":"5f7fbd813e94f16a85448745","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1649681653581-5f7fbd813e94f16a85448745.jpeg","fullname":"Sayak Paul","name":"sayakpaul","type":"user","isPro":true,"isHf":true,"isHfAdmin":false,"isMod":false,"followerCount":1032,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583856921041-5dd96eb166059660ed1ee413.png","fullname":"Hugging Face","name":"huggingface","type":"org","isHf":true,"details":"The AI community building the future.","plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9506018161773682},"editors":["sayakpaul"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1649681653581-5f7fbd813e94f16a85448745.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.27345","authors":[{"_id":"6a90f4fda64059bab69c3621","name":"Yuandong Pu","hidden":false},{"_id":"6a90f4fda64059bab69c3622","name":"Le Zhuo","hidden":false},{"_id":"6a90f4fda64059bab69c3623","user":{"_id":"5f7fbd813e94f16a85448745","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1649681653581-5f7fbd813e94f16a85448745.jpeg","isPro":true,"fullname":"Sayak Paul","user":"sayakpaul","type":"user","name":"sayakpaul"},"name":"Sayak Paul","status":"claimed_verified","statusLastChangedAt":"2026-08-28T08:45:04.630Z","hidden":false},{"_id":"6a90f4fda64059bab69c3624","name":"Gabriel Jorge Menezes","hidden":false},{"_id":"6a90f4fda64059bab69c3625","name":"Avram Đorđević","hidden":false},{"_id":"6a90f4fda64059bab69c3626","name":"Shiyang Li","hidden":false},{"_id":"6a90f4fda64059bab69c3627","name":"Yifan Zhou","hidden":false},{"_id":"6a90f4fda64059bab69c3628","name":"Bin Fu","hidden":false},{"_id":"6a90f4fda64059bab69c3629","name":"Wenlong Zhang","hidden":false},{"_id":"6a90f4fda64059bab69c362a","name":"Junjun He","hidden":false},{"_id":"6a90f4fda64059bab69c362b","name":"Yu Qiao","hidden":false},{"_id":"6a90f4fda64059bab69c362c","name":"Yihao Liu","hidden":false},{"_id":"6a90f4fda64059bab69c362d","name":"Jingbo Xing","hidden":false},{"_id":"6a90f4fda64059bab69c362e","name":"Xi Chen","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/5f7fbd813e94f16a85448745/ulHGPCQR4qq_tgEKK8FVE.png"],"publishedAt":"2026-08-27T00:00:00.000Z","submittedOnDailyAt":"2026-08-28T00:00:00.000Z","title":"PAWBench: How Far Are We from Probabilistically Aligned World Modeling?","submittedOnDailyBy":{"_id":"5f7fbd813e94f16a85448745","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1649681653581-5f7fbd813e94f16a85448745.jpeg","isPro":true,"fullname":"Sayak Paul","user":"sayakpaul","type":"user","name":"sayakpaul"},"summary":"Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.","upvotes":69,"discussionId":"6a90f4fda64059bab69c362f","projectPage":"https://pawbench.github.io/","githubRepo":"https://github.com/Andrew0613/PAWBench","githubRepoAddedBy":"user","ai_summary":"The study formalizes probabilistic alignment for world models, introduces PAWBench and PAWEval to evaluate video generators as stochastic samplers, and finds current models fail to match reference behavior distributions.","ai_keywords":["probabilistic alignment","world models","video generators","stochastic samplers","PAWBench","PAWEval","distributional criterion"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"625d5b9f0bec31f086e04cd9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1650285458447-noauth.jpeg","isPro":false,"fullname":"YuandongPu","user":"Andrew613","type":"user"},{"_id":"6a8d9cdec952a977218c1383","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8d9cdec952a977218c1383/dCH9ElvuQ3dWiguoHYw8n.jpeg","isPro":false,"fullname":"Sungmin HWANG","user":"sghwang84","type":"user"},{"_id":"659d2dff20cf0b934bbee513","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/659d2dff20cf0b934bbee513/9e9R852Zr2R82h64eUUQl.jpeg","isPro":false,"fullname":"Yifan Zhou","user":"yingmanji","type":"user"},{"_id":"6a8d9f82d68ef6b793a4a334","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8d9f82d68ef6b793a4a334/97PRpEQ5PTrqvEluMQ-ZY.jpeg","isPro":false,"fullname":"Emma E. Harris","user":"eharris02","type":"user"},{"_id":"662885b9b87483ae5a9ee5c9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/662885b9b87483ae5a9ee5c9/fLzCrWizo_hy7zGuAqFwk.jpeg","isPro":false,"fullname":"Songhao Han","user":"hshjerry0315","type":"user"},{"_id":"63e5adeb46965e1161d54ad4","avatarUrl":"/avatars/65ce4eabd26601d669688f9cca646893.svg","isPro":true,"fullname":"Liangbing Zhao","user":"metazlb","type":"user"},{"_id":"6434226da4c9c55871a78052","avatarUrl":"/avatars/3309832b3115bc6ad08ae1d10f43118b.svg","isPro":false,"fullname":"BoYang Zheng","user":"bytetriper","type":"user"},{"_id":"6a853fef9cce93fc04edd1bf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a853fef9cce93fc04edd1bf/_VyOkYd9RxuGxl6MM8kxZ.jpeg","isPro":false,"fullname":"Ayaan Kumar","user":"ayaankumarberg","type":"user"},{"_id":"6a815facd45e5232171fbdac","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a815facd45e5232171fbdac/yKZX8CGcXyHeICz1QkIYY.jpeg","isPro":false,"fullname":"राहुल कुमार","user":"Manuelrsgb2007","type":"user"},{"_id":"6a7f1bc6db98a5b9cbb64448","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a7f1bc6db98a5b9cbb64448/uVOBWTcw34tFchiW_vIM7.jpeg","isPro":false,"fullname":"Aarav","user":"AaravAgarwalna","type":"user"},{"_id":"6a85d6210037620809ff71b8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a85d6210037620809ff71b8/Vk0Ap3mKdwP7TcWA1t7qV.jpeg","isPro":false,"fullname":"Xin F. He","user":"xinhe3986","type":"user"},{"_id":"6358a167f56b03ec9147074d","avatarUrl":"/avatars/e54ea7bf0c240cf76d538296efb3976c.svg","isPro":false,"fullname":"Le Zhuo","user":"JackyZhuo","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.27345.md","query":{}}">
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
Abstract
The study formalizes probabilistic alignment for world models, introduces PAWBench and PAWEval to evaluate video generators as stochastic samplers, and finds current models fail to match reference behavior distributions.
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.
Community
We investigate the probabilistic failures in video generation models and present a reliable benchmark to measure it.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.27345 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.27345 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.