Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, yet must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid advances in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across isolated devices, tasks, and benchmarks. The central challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop. This survey is the first to systematically study smart glasses and develops a unified framework for investigating this loop. We formalize smart glasses through first-person data flow and constrained task utility, consolidate devices into eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks to datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. The resulting framework turns smart glasses into comparable, deployable, and reproducibly evaluated research objects, and outlines a roadmap toward trustworthy first-person intelligence, providing guidance for future research and applications in this rapidly evolving field.<br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/69de1d68da6d3334acc80ac5/93dChY4qmfDtZvf_yaSNu.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/69de1d68da6d3334acc80ac5/93dChY4qmfDtZvf_yaSNu.png\" alt=\"survey-pipeline\"></a></p>\n","updatedAt":"2026-08-26T03:36:46.233Z","author":{"_id":"69de1d68da6d3334acc80ac5","avatarUrl":"/avatars/77328d595f654df06be447c7a487f631.svg","fullname":"Haojun Chen","name":"HaojunChen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.863526463508606},"editors":["HaojunChen"],"editorAvatarUrls":["/avatars/77328d595f654df06be447c7a487f631.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.24877","authors":[{"_id":"6a8e4a397bc881afa25f306f","name":"Jiangning Zhang","hidden":false},{"_id":"6a8e4a397bc881afa25f3070","name":"Haojun Chen","hidden":false},{"_id":"6a8e4a397bc881afa25f3071","name":"Yong Liu","hidden":false}],"publishedAt":"2026-08-25T00:00:00.000Z","submittedOnDailyAt":"2026-08-26T00:00:00.000Z","title":"From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms","submittedOnDailyBy":{"_id":"69de1d68da6d3334acc80ac5","avatarUrl":"/avatars/77328d595f654df06be447c7a487f631.svg","isPro":false,"fullname":"Haojun Chen","user":"HaojunChen","type":"user","name":"HaojunChen"},"summary":"Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop. This survey is the \\textbf{first to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.","upvotes":8,"discussionId":"6a8e4a397bc881afa25f3072","projectPage":"https://github.com/zhangzjn/awesome-smart-glasses","githubRepo":"https://github.com/zhangzjn/awesome-smart-glasses","githubRepoAddedBy":"user","ai_summary":"Smart glasses are surveyed as unified first-person intelligence platforms requiring sustained perception-state-interaction-action loops across constrained hardware and diverse applications.","ai_keywords":["multimodal models","embodied intelligence","egocentric vision"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"686100b5242fe044d66171e2","avatarUrl":"/avatars/757c93b4878e5bb55faa488a0209ee73.svg","isPro":false,"fullname":"Haojun Chen","user":"HouzeonZan","type":"user"},{"_id":"69de1d68da6d3334acc80ac5","avatarUrl":"/avatars/77328d595f654df06be447c7a487f631.svg","isPro":false,"fullname":"Haojun Chen","user":"HaojunChen","type":"user"},{"_id":"63fa1f88d38275b44359398d","avatarUrl":"/avatars/af7a826e59263dd8272368927a6930fb.svg","isPro":false,"fullname":"APRIL-AIGC","user":"APRIL-AIGC","type":"user"},{"_id":"68b7bb40f02488c467e3d54b","avatarUrl":"/avatars/d325ff6c3edc075c347196242b3f273f.svg","isPro":false,"fullname":"Eddie0521","user":"Eddie0521","type":"user"},{"_id":"6691f7463e80c7fa549c1092","avatarUrl":"/avatars/a02b7b9722f6554695f4f210872c86f1.svg","isPro":false,"fullname":"Yinan Chen","user":"Coraxor","type":"user"},{"_id":"64d761b98ebc40443831f82a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64d761b98ebc40443831f82a/DHBOtOstiFp2-lDY6b9gb.png","isPro":false,"fullname":"Guangyi Liu","user":"lgy0404","type":"user"},{"_id":"65755c08b238c76bbaed1c4c","avatarUrl":"/avatars/17790ad1a0940ad8450b7a967133bce8.svg","isPro":false,"fullname":"WangYuji","user":"YujiWang132","type":"user"},{"_id":"61af81009f77f7b669578f95","avatarUrl":"/avatars/fb50773ac49948940eb231834ee6f2fd.svg","isPro":false,"fullname":"rotem israeli","user":"irotem98","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.24877.md","query":{}}">
From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
Abstract
Smart glasses are surveyed as unified first-person intelligence platforms requiring sustained perception-state-interaction-action loops across constrained hardware and diverse applications.
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop. This survey is the \textbf{first to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.
Community
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, yet must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid advances in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across isolated devices, tasks, and benchmarks. The central challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop. This survey is the first to systematically study smart glasses and develops a unified framework for investigating this loop. We formalize smart glasses through first-person data flow and constrained task utility, consolidate devices into eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks to datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. The resulting framework turns smart glasses into comparable, deployable, and reproducibly evaluated research objects, and outlines a roadmap toward trustworthy first-person intelligence, providing guidance for future research and applications in this rapidly evolving field.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.24877 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.24877 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.24877 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.