We propose AnTrap, a dynamic adversarial evaluation framework that injects realistic runtime anomalies across State, Thinking, Action, and Round levels to stress-test Android GUI agents, uncovering universal performance degradation and intrinsic reasoning limitations under deep contextual traps.</p>\n","updatedAt":"2026-08-27T02:44:52.056Z","author":{"_id":"65dfeee3d16fb170031df293","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65dfeee3d16fb170031df293/2VbNuqcpN3XrWB18NfzRQ.jpeg","fullname":"gan","name":"guo9","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8201957941055298},"editors":["guo9"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/65dfeee3d16fb170031df293/2VbNuqcpN3XrWB18NfzRQ.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.24099","authors":[{"_id":"6a8fa3402c24e8c5fab3295d","name":"Guo Gan","hidden":false},{"_id":"6a8fa3402c24e8c5fab3295e","name":"Yilun Zhao","hidden":false},{"_id":"6a8fa3402c24e8c5fab3295f","name":"Cong Chen","hidden":false},{"_id":"6a8fa3402c24e8c5fab32960","name":"Jinbiao Wei","hidden":false},{"_id":"6a8fa3402c24e8c5fab32961","name":"Tingyu Song","hidden":false},{"_id":"6a8fa3402c24e8c5fab32962","name":"Zheyuan Yang","hidden":false},{"_id":"6a8fa3402c24e8c5fab32963","name":"Lin Fu","hidden":false},{"_id":"6a8fa3402c24e8c5fab32964","name":"Hong Zhou","hidden":false}],"publishedAt":"2026-08-25T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments","submittedOnDailyBy":{"_id":"65dfeee3d16fb170031df293","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65dfeee3d16fb170031df293/2VbNuqcpN3XrWB18NfzRQ.jpeg","isPro":false,"fullname":"gan","user":"guo9","type":"user","name":"guo9"},"summary":"GUI agents often encounter dynamic anomalies when deployed on Android devices, from unexpected pop-ups to action misuse, yet existing benchmarks lack systematic evaluation of agent robustness against runtime anomalies. We introduce AnTrap, a comprehensive benchmark that injects dynamic perturbations into agent execution trajectories. We propose a taxonomy organizing real-world anomalies into four layers (State, Thinking, Action and Round) with ten fine-grained subcategories, and develop a construction pipeline that preserves task solvability while introducing realistic adversarial conditions. Evaluating 16 leading GUI models, we reveal universal vulnerability to dynamic anomalies, with even the strongest models suffering significant performance degradation. Furthermore, we conduct GRPO training in both original and adversarial environments to validate our benchmark, separating environment-learnable anomalies from reasoning-bottlenecked ones. Our findings show that while single-step traps at state and action layers are largely addressable through adversarial reinforcement learning, deep contextual traps, like state deadlock, expose intrinsic limitations that cannot be resolved by training in environments with traps alone.","upvotes":12,"discussionId":"6a8fa3402c24e8c5fab32965","githubRepo":"https://github.com/gguogan/AnTrap","githubRepoAddedBy":"user","ai_summary":"AnTrap benchmarks GUI agent robustness by injecting dynamic anomalies into execution trajectories, revealing universal vulnerabilities and distinguishing learnable traps from intrinsic reasoning limits.","ai_keywords":["GUI agents","dynamic anomalies","AnTrap","taxonomy","adversarial conditions","GRPO training","reinforcement learning","state deadlock"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6549ab205018913069fb8eab","avatarUrl":"/avatars/30e09cda80a2bcca9100e3464c175529.svg","isPro":false,"fullname":"chencong","user":"Chencong1","type":"user"},{"_id":"65dfeee3d16fb170031df293","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65dfeee3d16fb170031df293/2VbNuqcpN3XrWB18NfzRQ.jpeg","isPro":false,"fullname":"gan","user":"guo9","type":"user"},{"_id":"62f662bcc58915315c4eccea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62f662bcc58915315c4eccea/zOAQLONfMP88zr70sxHK-.jpeg","isPro":true,"fullname":"Yilun Zhao","user":"yilunzhao","type":"user"},{"_id":"68084d54aca60e6178b3afb5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68084d54aca60e6178b3afb5/TshN3Ka3VRFD_I3WJ6Vys.jpeg","isPro":false,"fullname":"Lin Fu","user":"minuzero","type":"user"},{"_id":"66af69222f4c59963afc874f","avatarUrl":"/avatars/034ca7688282bdbeddbd4f03e54dead7.svg","isPro":false,"fullname":"Zheyuan Yang","user":"Raywithyou","type":"user"},{"_id":"63f58403fcf95ecac2b33d78","avatarUrl":"/avatars/a77ea80784896502ae1cfa086a78ce66.svg","isPro":false,"fullname":"Zhen Yang","user":"YZCS","type":"user"},{"_id":"6431974f034ecbefddd4b463","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/O6yDcmDTZk9IH2wIbuCQ3.jpeg","isPro":false,"fullname":"刘自得","user":"zideliu","type":"user"},{"_id":"63c9537586529da209591cf1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674138561431-63c9537586529da209591cf1.jpeg","isPro":false,"fullname":"Chengxiang Fan","user":"leaf1170124460","type":"user"},{"_id":"64dc29d9b5d625e0e9a6ecb9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QxGBsnk1cNsBEPqSx4ae-.jpeg","isPro":false,"fullname":"Tingyu Song","user":"songtingyu","type":"user"},{"_id":"69941106d3644717b90e08fc","avatarUrl":"/avatars/fe30ed4f241ee6dcac7d39b784413988.svg","isPro":false,"fullname":"EDIR","user":"EDIR-BENCH","type":"user"},{"_id":"644238f6fcbe90d73b319ea6","avatarUrl":"/avatars/06dcede0ea77344de7dded917706bcb2.svg","isPro":false,"fullname":"zhongzero","user":"zhongzero","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.24099.md","query":{}}">
Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments
Published on Aug 25
· Submitted by gan on Aug 27 Abstract
AnTrap benchmarks GUI agent robustness by injecting dynamic anomalies into execution trajectories, revealing universal vulnerabilities and distinguishing learnable traps from intrinsic reasoning limits.
GUI agents often encounter dynamic anomalies when deployed on Android devices, from unexpected pop-ups to action misuse, yet existing benchmarks lack systematic evaluation of agent robustness against runtime anomalies. We introduce AnTrap, a comprehensive benchmark that injects dynamic perturbations into agent execution trajectories. We propose a taxonomy organizing real-world anomalies into four layers (State, Thinking, Action and Round) with ten fine-grained subcategories, and develop a construction pipeline that preserves task solvability while introducing realistic adversarial conditions. Evaluating 16 leading GUI models, we reveal universal vulnerability to dynamic anomalies, with even the strongest models suffering significant performance degradation. Furthermore, we conduct GRPO training in both original and adversarial environments to validate our benchmark, separating environment-learnable anomalies from reasoning-bottlenecked ones. Our findings show that while single-step traps at state and action layers are largely addressable through adversarial reinforcement learning, deep contextual traps, like state deadlock, expose intrinsic limitations that cannot be resolved by training in environments with traps alone.
Community
We propose AnTrap, a dynamic adversarial evaluation framework that injects realistic runtime anomalies across State, Thinking, Action, and Round levels to stress-test Android GUI agents, uncovering universal performance degradation and intrinsic reasoning limitations under deep contextual traps.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.24099 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.24099 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.24099 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.