MobilePA-Bench</p>\n","updatedAt":"2026-08-25T05:58:44.066Z","author":{"_id":"6246bb33da617c00b48e4d92","avatarUrl":"/avatars/0304a9f6eb7f5dee4d933d03222f94e9.svg","fullname":"Weigao Sun","name":"weigao266","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5590533018112183},"editors":["weigao266"],"editorAvatarUrls":["/avatars/0304a9f6eb7f5dee4d933d03222f94e9.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.23035","authors":[{"_id":"6a8d24dc5add2537c32e97c6","name":"Yi Zhu","hidden":false},{"_id":"6a8d24dc5add2537c32e97c7","user":{"_id":"642521d1a4f3051f54dd2935","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642521d1a4f3051f54dd2935/8sxV_8qXVdajrrFpJNMgX.png","isPro":false,"fullname":"xwwu","user":"xwwu","type":"user","name":"xwwu"},"name":"Xiongwei Wu","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:13:36.898Z","hidden":false},{"_id":"6a8d24dc5add2537c32e97c8","name":"Qiyi Wang","hidden":false},{"_id":"6a8d24dc5add2537c32e97c9","user":{"_id":"64d9e5d276daedd6b1bd3155","avatarUrl":"/avatars/bfb235fb5e036a7f05e08d2f9781dff1.svg","isPro":false,"fullname":"Tingyu Qu","user":"tingyuqu95","type":"user","name":"tingyuqu95"},"name":"Tingyu Qu","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:13:34.856Z","hidden":false},{"_id":"6a8d24dc5add2537c32e97ca","name":"Jiajun Liu","hidden":false},{"_id":"6a8d24dc5add2537c32e97cb","name":"Sihan Cao","hidden":false},{"_id":"6a8d24dc5add2537c32e97cc","name":"Long Chen","hidden":false},{"_id":"6a8d24dc5add2537c32e97cd","user":{"_id":"6246bb33da617c00b48e4d92","avatarUrl":"/avatars/0304a9f6eb7f5dee4d933d03222f94e9.svg","isPro":false,"fullname":"Weigao Sun","user":"weigao266","type":"user","name":"weigao266"},"name":"Weigao Sun","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:13:33.097Z","hidden":false},{"_id":"6a8d24dc5add2537c32e97ce","name":"Feida Zhu","hidden":false},{"_id":"6a8d24dc5add2537c32e97cf","name":"Yiran Zhong","hidden":false},{"_id":"6a8d24dc5add2537c32e97d0","name":"Steven Hoi","hidden":false}],"publishedAt":"2026-08-24T00:00:00.000Z","submittedOnDailyAt":"2026-08-25T00:00:00.000Z","title":"MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks","submittedOnDailyBy":{"_id":"6246bb33da617c00b48e4d92","avatarUrl":"/avatars/0304a9f6eb7f5dee4d933d03222f94e9.svg","isPro":false,"fullname":"Weigao Sun","user":"weigao266","type":"user","name":"weigao266"},"summary":"As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps, each with a critical blind spot: GUI-centric benchmarks test surface-level screen manipulation while overlooking background tool use and long-horizon planning, whereas static function-calling benchmarks rely on offline API matching that is detached from real runtime constraints. To close this gap, we present MobilePA-Bench, an interactive, stateful, and tool-centric benchmark for evaluating the tool-calling and planning abilities of mobile planning agents. MobilePA-Bench runs on an executable sandbox that maintains live application databases and returns structured feedback, spanning 13 functional domains and 212 realistic mobile tools. Beyond basic tool use, it evaluates a central planning agent along three advanced dimensions: (1)~Sub-agent Collaboration---decomposing a complex task and delegating specialized work to capable sub-agents; (2)~Memory Usage---recalling stored memories, user profiles, and past preferences to resolve implicit requests; and (3)~Skill Usage---invoking pre-packaged composite skills instead of planning every step from scratch. Extensive experiments show that current frontier LLMs remain unreliable in mobile settings: performance drops sharply under strict tool ordering, permission limits, and unexpected runtime errors. By pairing an interactive function-calling sandbox with evidence-based verification, MobilePA-Bench serves as both a practical diagnostic benchmark and an interactive foundation for agentic reinforcement learning---accelerating the development of dependable mobile agents.","upvotes":31,"discussionId":"6a8d24dc5add2537c32e97d1","projectPage":"https://tongyi-mai.github.io/MobilePA-Bench/","ai_summary":"MobilePA-Bench is an interactive sandbox benchmark that evaluates mobile planning agents on tool-calling, sub-agent collaboration, memory usage, and composite skill invocation under real runtime constraints.","ai_keywords":["tool-calling","mobile planning agents","executable sandbox","sub-agent collaboration","memory usage","skill usage","composite skills","function-calling sandbox","reinforcement learning","LLMs"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"6925b20fed452d1567c012d3","name":"Tongyi-MAI","fullname":"Tongyi-MAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64379d79fac5ea753f1c10f3/fxHO6QoYjdv9_LTyiUD3g.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"642521d1a4f3051f54dd2935","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642521d1a4f3051f54dd2935/8sxV_8qXVdajrrFpJNMgX.png","isPro":false,"fullname":"xwwu","user":"xwwu","type":"user"},{"_id":"65a53fbcc8a09bd5e84873e7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a53fbcc8a09bd5e84873e7/dgt6YdgugevSR0R-jyEvT.jpeg","isPro":false,"fullname":"yuhaowang","user":"NukaSarsae","type":"user"},{"_id":"67777b7a8376dfe003afa951","avatarUrl":"/avatars/2af9d3181306d4c53329d047eeadaf1e.svg","isPro":false,"fullname":"Sihan Cao","user":"Sihan-Cao","type":"user"},{"_id":"64d9e5d276daedd6b1bd3155","avatarUrl":"/avatars/bfb235fb5e036a7f05e08d2f9781dff1.svg","isPro":false,"fullname":"Tingyu Qu","user":"tingyuqu95","type":"user"},{"_id":"64f6efa0e2e54c750fb0655d","avatarUrl":"/avatars/83518aff2e2ef0f9cf84ca39e0e5db0d.svg","isPro":false,"fullname":"Ming Ma","user":"MBJinX","type":"user"},{"_id":"645c6b8f4784d6388415d459","avatarUrl":"/avatars/91ae69959dba89be41259f678ca2fc21.svg","isPro":false,"fullname":"Zhu","user":"Jason0102","type":"user"},{"_id":"646def60df618b303b419323","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646def60df618b303b419323/JLJGYen4-5M8ivsLsSk0w.jpeg","isPro":false,"fullname":"Lei Wang","user":"demolei","type":"user"},{"_id":"6a2c2f829c1cedacc47787f3","avatarUrl":"/avatars/7e446c3cb0a8a4918b0bc822d8267654.svg","isPro":false,"fullname":"Wang","user":"Qichao25","type":"user"},{"_id":"6246bb33da617c00b48e4d92","avatarUrl":"/avatars/0304a9f6eb7f5dee4d933d03222f94e9.svg","isPro":false,"fullname":"Weigao Sun","user":"weigao266","type":"user"},{"_id":"63fc8055d44f50f5595884c3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/NiQJgxSYQkeTKABQ4yJcd.png","isPro":false,"fullname":"Jianqiang Ren","user":"QiangBro","type":"user"},{"_id":"668aa356c6f9383fb758cd66","avatarUrl":"/avatars/38bc9ae61a8f4f683d9a64fc616fd20b.svg","isPro":false,"fullname":"hhh","user":"MXwanghhh","type":"user"},{"_id":"64cb238576200ec80fe988f8","avatarUrl":"/avatars/42c48710c7881c9dfbcc075fec3cb600.svg","isPro":false,"fullname":"zeus","user":"zengw","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6925b20fed452d1567c012d3","name":"Tongyi-MAI","fullname":"Tongyi-MAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64379d79fac5ea753f1c10f3/fxHO6QoYjdv9_LTyiUD3g.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.23035.md","query":{}}">
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
Abstract
MobilePA-Bench is an interactive sandbox benchmark that evaluates mobile planning agents on tool-calling, sub-agent collaboration, memory usage, and composite skill invocation under real runtime constraints.
As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps, each with a critical blind spot: GUI-centric benchmarks test surface-level screen manipulation while overlooking background tool use and long-horizon planning, whereas static function-calling benchmarks rely on offline API matching that is detached from real runtime constraints. To close this gap, we present MobilePA-Bench, an interactive, stateful, and tool-centric benchmark for evaluating the tool-calling and planning abilities of mobile planning agents. MobilePA-Bench runs on an executable sandbox that maintains live application databases and returns structured feedback, spanning 13 functional domains and 212 realistic mobile tools. Beyond basic tool use, it evaluates a central planning agent along three advanced dimensions: (1)~Sub-agent Collaboration---decomposing a complex task and delegating specialized work to capable sub-agents; (2)~Memory Usage---recalling stored memories, user profiles, and past preferences to resolve implicit requests; and (3)~Skill Usage---invoking pre-packaged composite skills instead of planning every step from scratch. Extensive experiments show that current frontier LLMs remain unreliable in mobile settings: performance drops sharply under strict tool ordering, permission limits, and unexpected runtime errors. By pairing an interactive function-calling sandbox with evidence-based verification, MobilePA-Bench serves as both a practical diagnostic benchmark and an interactive foundation for agentic reinforcement learning---accelerating the development of dependable mobile agents.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.23035 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.23035 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.23035 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.