Hugging Face Daily Papers · · 4 min read

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots and environment interactive signals. Authors are from Tsinghua University and Tencent Hunyuan. </p>\n","updatedAt":"2026-06-30T02:25:34.970Z","author":{"_id":"6434c9dc4b34368fdb07d421","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6434c9dc4b34368fdb07d421/V_afg81iuNyMFfhM7qdgB.jpeg","fullname":"fansunqi","name":"fansunqi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.7427505254745483},"editors":["fansunqi"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6434c9dc4b34368fdb07d421/V_afg81iuNyMFfhM7qdgB.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.29705","authors":[{"_id":"6a432851763f63ca3757e827","name":"Sunqi Fan","hidden":false},{"_id":"6a432851763f63ca3757e828","name":"Lingshan Chen","hidden":false},{"_id":"6a432851763f63ca3757e829","name":"Runqi Yin","hidden":false},{"_id":"6a432851763f63ca3757e82a","name":"Qingle Liu","hidden":false},{"_id":"6a432851763f63ca3757e82b","name":"Yongming Rao","hidden":false},{"_id":"6a432851763f63ca3757e82c","name":"Meng-Hao Guo","hidden":false},{"_id":"6a432851763f63ca3757e82d","name":"Shi-Min Hu","hidden":false}],"publishedAt":"2026-06-29T00:00:00.000Z","submittedOnDailyAt":"2026-06-30T00:00:00.000Z","title":"GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots","submittedOnDailyBy":{"_id":"6434c9dc4b34368fdb07d421","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6434c9dc4b34368fdb07d421/V_afg81iuNyMFfhM7qdgB.jpeg","isPro":false,"fullname":"fansunqi","user":"fansunqi","type":"user","name":"fansunqi"},"summary":"Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models. Naturally, researchers aim to extend this paradigm to the domain of GUI agents, hoping to build strong GUI agents through a similar paradigm. However, GUI agent data cannot be directly harvested from the internet, making it costly and difficult to collect at scale. As a result, current GUI agents suffer from poor cross-device generalization and limited visual grounding ability for fine-grained GUI elements. As an attempt to address data challenge in GUI agents, we propose GUICrafter, a weakly-supervised GUI agent leveraging massive unannotated screenshots to substantially reduce the reliance on expensive human annotations. GUICrafter explores a curriculum learning framework for training GUI agents through two progressive stages. First, the model learns visual grounding from large-scale unannotated screenshots and webpages, leveraging the rich contextual signals inherent in GUI interactions without human annotations. Then, in Stage 2, we leverage a small amount of high-quality data to calibrate the model via reinforcement learning. Experiments show that GUICrafter achieves competitive, or even superior, performance to advanced systems like UI-TARS while using only 0.1% of its data. Furthermore, under the same amount of annotated data, GUICrafter surpasses all previous methods such as GUI-R1. Code, data, and models are available at https://github.com/fansunqi/GUICrafter.","upvotes":10,"discussionId":"6a432851763f63ca3757e82e","githubRepo":"https://github.com/fansunqi/GUICrafter","githubRepoAddedBy":"user","ai_summary":"GUICrafter addresses GUI agent data challenges through a weakly-supervised approach using unannotated screenshots and a two-stage curriculum learning framework for visual grounding and reinforcement learning calibration.","ai_keywords":["GUI agents","weakly-supervised learning","curriculum learning","visual grounding","reinforcement learning","GUI interaction","screen shots","webpages"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":2,"organization":{"_id":"6645f953c39288df638dbdd5","name":"Tencent-Hunyuan","fullname":"Tencent Hunyuan","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d22496c58f969c152bcefd/woKSjt2wXvBNKussyYPsa.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6434c9dc4b34368fdb07d421","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6434c9dc4b34368fdb07d421/V_afg81iuNyMFfhM7qdgB.jpeg","isPro":false,"fullname":"fansunqi","user":"fansunqi","type":"user"},{"_id":"69d759ecfa5f735977291448","avatarUrl":"/avatars/7bc512e8cc06505c091533e68daed956.svg","isPro":false,"fullname":"Stephen Fan","user":"stephenfan1101","type":"user"},{"_id":"6869ecb0118df64ef629b43d","avatarUrl":"/avatars/a0c973ed7b1b83ade1a8569ca8747c55.svg","isPro":false,"fullname":"Lingshan Chen","user":"chen03","type":"user"},{"_id":"66f0d2036a483077eed42bfb","avatarUrl":"/avatars/f7f3f726842c26b8e52c9bdd48774b8e.svg","isPro":false,"fullname":"Renping Zhou","user":"rpzhou","type":"user"},{"_id":"6909c3259ceb1a82bd511c08","avatarUrl":"/avatars/9286e0b9231751277c14e1fdc6e78fe7.svg","isPro":false,"fullname":"Runqi Yin","user":"yinrunqi","type":"user"},{"_id":"64ac0605091ff88865352b44","avatarUrl":"/avatars/c28acb08a1fdeab899afd1961d3dd94a.svg","isPro":false,"fullname":"ytz","user":"ytz20","type":"user"},{"_id":"67cc5e0232aeea9209d35033","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/bgyP6thb5Wb4cUENl6cDF.png","isPro":false,"fullname":"YueChen","user":"YueChen0614","type":"user"},{"_id":"68fdf676e39d2cbbb5a4316f","avatarUrl":"/avatars/6a8f7a2e0178e6aba2d3c9fa8181462c.svg","isPro":false,"fullname":"Qingle Liu","user":"Aoraku","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6645f953c39288df638dbdd5","name":"Tencent-Hunyuan","fullname":"Tencent Hunyuan","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d22496c58f969c152bcefd/woKSjt2wXvBNKussyYPsa.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.29705.md","query":{}}">
Papers
arxiv:2606.29705

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots

Published on Jun 29
· Submitted by
fansunqi
on Jun 30
Authors:
,
,
,
,
,
,

Abstract

GUICrafter addresses GUI agent data challenges through a weakly-supervised approach using unannotated screenshots and a two-stage curriculum learning framework for visual grounding and reinforcement learning calibration.

Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models. Naturally, researchers aim to extend this paradigm to the domain of GUI agents, hoping to build strong GUI agents through a similar paradigm. However, GUI agent data cannot be directly harvested from the internet, making it costly and difficult to collect at scale. As a result, current GUI agents suffer from poor cross-device generalization and limited visual grounding ability for fine-grained GUI elements. As an attempt to address data challenge in GUI agents, we propose GUICrafter, a weakly-supervised GUI agent leveraging massive unannotated screenshots to substantially reduce the reliance on expensive human annotations. GUICrafter explores a curriculum learning framework for training GUI agents through two progressive stages. First, the model learns visual grounding from large-scale unannotated screenshots and webpages, leveraging the rich contextual signals inherent in GUI interactions without human annotations. Then, in Stage 2, we leverage a small amount of high-quality data to calibrate the model via reinforcement learning. Experiments show that GUICrafter achieves competitive, or even superior, performance to advanced systems like UI-TARS while using only 0.1% of its data. Furthermore, under the same amount of annotated data, GUICrafter surpasses all previous methods such as GUI-R1. Code, data, and models are available at https://github.com/fansunqi/GUICrafter.

Community

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots and environment interactive signals. Authors are from Tsinghua University and Tencent Hunyuan.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.29705
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2606.29705 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2606.29705 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.29705 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers