Hugging Face Daily Papers · · 4 min read

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

MobilePA-Bench</p>\n","updatedAt":"2026-08-25T05:58:44.066Z","author":{"_id":"6246bb33da617c00b48e4d92","avatarUrl":"/avatars/0304a9f6eb7f5dee4d933d03222f94e9.svg","fullname":"Weigao Sun","name":"weigao266","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5590533018112183},"editors":["weigao266"],"editorAvatarUrls":["/avatars/0304a9f6eb7f5dee4d933d03222f94e9.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.23035","authors":[{"_id":"6a8d24dc5add2537c32e97c6","name":"Yi Zhu","hidden":false},{"_id":"6a8d24dc5add2537c32e97c7","user":{"_id":"642521d1a4f3051f54dd2935","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642521d1a4f3051f54dd2935/8sxV_8qXVdajrrFpJNMgX.png","isPro":false,"fullname":"xwwu","user":"xwwu","type":"user","name":"xwwu"},"name":"Xiongwei Wu","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:13:36.898Z","hidden":false},{"_id":"6a8d24dc5add2537c32e97c8","name":"Qiyi Wang","hidden":false},{"_id":"6a8d24dc5add2537c32e97c9","user":{"_id":"64d9e5d276daedd6b1bd3155","avatarUrl":"/avatars/bfb235fb5e036a7f05e08d2f9781dff1.svg","isPro":false,"fullname":"Tingyu Qu","user":"tingyuqu95","type":"user","name":"tingyuqu95"},"name":"Tingyu Qu","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:13:34.856Z","hidden":false},{"_id":"6a8d24dc5add2537c32e97ca","name":"Jiajun Liu","hidden":false},{"_id":"6a8d24dc5add2537c32e97cb","name":"Sihan Cao","hidden":false},{"_id":"6a8d24dc5add2537c32e97cc","name":"Long Chen","hidden":false},{"_id":"6a8d24dc5add2537c32e97cd","user":{"_id":"6246bb33da617c00b48e4d92","avatarUrl":"/avatars/0304a9f6eb7f5dee4d933d03222f94e9.svg","isPro":false,"fullname":"Weigao Sun","user":"weigao266","type":"user","name":"weigao266"},"name":"Weigao Sun","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:13:33.097Z","hidden":false},{"_id":"6a8d24dc5add2537c32e97ce","name":"Feida Zhu","hidden":false},{"_id":"6a8d24dc5add2537c32e97cf","name":"Yiran Zhong","hidden":false},{"_id":"6a8d24dc5add2537c32e97d0","name":"Steven Hoi","hidden":false}],"publishedAt":"2026-08-24T00:00:00.000Z","submittedOnDailyAt":"2026-08-25T00:00:00.000Z","title":"MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks","submittedOnDailyBy":{"_id":"6246bb33da617c00b48e4d92","avatarUrl":"/avatars/0304a9f6eb7f5dee4d933d03222f94e9.svg","isPro":false,"fullname":"Weigao Sun","user":"weigao266","type":"user","name":"weigao266"},"summary":"As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps, each with a critical blind spot: GUI-centric benchmarks test surface-level screen manipulation while overlooking background tool use and long-horizon planning, whereas static function-calling benchmarks rely on offline API matching that is detached from real runtime constraints. To close this gap, we present MobilePA-Bench, an interactive, stateful, and tool-centric benchmark for evaluating the tool-calling and planning abilities of mobile planning agents. MobilePA-Bench runs on an executable sandbox that maintains live application databases and returns structured feedback, spanning 13 functional domains and 212 realistic mobile tools. Beyond basic tool use, it evaluates a central planning agent along three advanced dimensions: (1)~Sub-agent Collaboration---decomposing a complex task and delegating specialized work to capable sub-agents; (2)~Memory Usage---recalling stored memories, user profiles, and past preferences to resolve implicit requests; and (3)~Skill Usage---invoking pre-packaged composite skills instead of planning every step from scratch. Extensive experiments show that current frontier LLMs remain unreliable in mobile settings: performance drops sharply under strict tool ordering, permission limits, and unexpected runtime errors. By pairing an interactive function-calling sandbox with evidence-based verification, MobilePA-Bench serves as both a practical diagnostic benchmark and an interactive foundation for agentic reinforcement learning---accelerating the development of dependable mobile agents.","upvotes":31,"discussionId":"6a8d24dc5add2537c32e97d1","projectPage":"https://tongyi-mai.github.io/MobilePA-Bench/","ai_summary":"MobilePA-Bench is an interactive sandbox benchmark that evaluates mobile planning agents on tool-calling, sub-agent collaboration, memory usage, and composite skill invocation under real runtime constraints.","ai_keywords":["tool-calling","mobile planning agents","executable sandbox","sub-agent collaboration","memory usage","skill usage","composite skills","function-calling sandbox","reinforcement learning","LLMs"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"6925b20fed452d1567c012d3","name":"Tongyi-MAI","fullname":"Tongyi-MAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64379d79fac5ea753f1c10f3/fxHO6QoYjdv9_LTyiUD3g.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"642521d1a4f3051f54dd2935","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642521d1a4f3051f54dd2935/8sxV_8qXVdajrrFpJNMgX.png","isPro":false,"fullname":"xwwu","user":"xwwu","type":"user"},{"_id":"65a53fbcc8a09bd5e84873e7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a53fbcc8a09bd5e84873e7/dgt6YdgugevSR0R-jyEvT.jpeg","isPro":false,"fullname":"yuhaowang","user":"NukaSarsae","type":"user"},{"_id":"67777b7a8376dfe003afa951","avatarUrl":"/avatars/2af9d3181306d4c53329d047eeadaf1e.svg","isPro":false,"fullname":"Sihan Cao","user":"Sihan-Cao","type":"user"},{"_id":"64d9e5d276daedd6b1bd3155","avatarUrl":"/avatars/bfb235fb5e036a7f05e08d2f9781dff1.svg","isPro":false,"fullname":"Tingyu Qu","user":"tingyuqu95","type":"user"},{"_id":"64f6efa0e2e54c750fb0655d","avatarUrl":"/avatars/83518aff2e2ef0f9cf84ca39e0e5db0d.svg","isPro":false,"fullname":"Ming Ma","user":"MBJinX","type":"user"},{"_id":"645c6b8f4784d6388415d459","avatarUrl":"/avatars/91ae69959dba89be41259f678ca2fc21.svg","isPro":false,"fullname":"Zhu","user":"Jason0102","type":"user"},{"_id":"646def60df618b303b419323","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646def60df618b303b419323/JLJGYen4-5M8ivsLsSk0w.jpeg","isPro":false,"fullname":"Lei Wang","user":"demolei","type":"user"},{"_id":"6a2c2f829c1cedacc47787f3","avatarUrl":"/avatars/7e446c3cb0a8a4918b0bc822d8267654.svg","isPro":false,"fullname":"Wang","user":"Qichao25","type":"user"},{"_id":"6246bb33da617c00b48e4d92","avatarUrl":"/avatars/0304a9f6eb7f5dee4d933d03222f94e9.svg","isPro":false,"fullname":"Weigao Sun","user":"weigao266","type":"user"},{"_id":"63fc8055d44f50f5595884c3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/NiQJgxSYQkeTKABQ4yJcd.png","isPro":false,"fullname":"Jianqiang Ren","user":"QiangBro","type":"user"},{"_id":"668aa356c6f9383fb758cd66","avatarUrl":"/avatars/38bc9ae61a8f4f683d9a64fc616fd20b.svg","isPro":false,"fullname":"hhh","user":"MXwanghhh","type":"user"},{"_id":"64cb238576200ec80fe988f8","avatarUrl":"/avatars/42c48710c7881c9dfbcc075fec3cb600.svg","isPro":false,"fullname":"zeus","user":"zengw","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6925b20fed452d1567c012d3","name":"Tongyi-MAI","fullname":"Tongyi-MAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64379d79fac5ea753f1c10f3/fxHO6QoYjdv9_LTyiUD3g.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.23035.md","query":{}}">
Papers
arxiv:2608.23035

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

Published on Aug 24
· Submitted by
Weigao Sun
on Aug 25
Authors:
,

Abstract

MobilePA-Bench is an interactive sandbox benchmark that evaluates mobile planning agents on tool-calling, sub-agent collaboration, memory usage, and composite skill invocation under real runtime constraints.

As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps, each with a critical blind spot: GUI-centric benchmarks test surface-level screen manipulation while overlooking background tool use and long-horizon planning, whereas static function-calling benchmarks rely on offline API matching that is detached from real runtime constraints. To close this gap, we present MobilePA-Bench, an interactive, stateful, and tool-centric benchmark for evaluating the tool-calling and planning abilities of mobile planning agents. MobilePA-Bench runs on an executable sandbox that maintains live application databases and returns structured feedback, spanning 13 functional domains and 212 realistic mobile tools. Beyond basic tool use, it evaluates a central planning agent along three advanced dimensions: (1)~Sub-agent Collaboration---decomposing a complex task and delegating specialized work to capable sub-agents; (2)~Memory Usage---recalling stored memories, user profiles, and past preferences to resolve implicit requests; and (3)~Skill Usage---invoking pre-packaged composite skills instead of planning every step from scratch. Extensive experiments show that current frontier LLMs remain unreliable in mobile settings: performance drops sharply under strict tool ordering, permission limits, and unexpected runtime errors. By pairing an interactive function-calling sandbox with evidence-based verification, MobilePA-Bench serves as both a practical diagnostic benchmark and an interactive foundation for agentic reinforcement learning---accelerating the development of dependable mobile agents.

Community

Paper author Paper submitter about 3 hours ago

MobilePA-Bench

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.23035
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.23035 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.23035 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.23035 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers