從看見到行動:智慧眼鏡作為第一人稱智慧平台
From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
August 25, 2026
作者: Jiangning Zhang, Haojun Chen, Yong Liu
cs.AI
摘要
智慧眼鏡正從拍攝與顯示配件,演進為連結人類感知、持續情境脈絡與數位或實體行動的第一人稱智能平台。其穿戴式視角與佩戴者的視覺、聽覺、動作及手部-物體互動對齊,但必須在嚴格的能耗、散熱、隱私與回饋限制下運作。儘管擴增實境、自我中心視覺、多模態模型、人機互動與具身智能進展迅速,文獻在裝置、任務與基準測試之間仍顯破碎。關鍵挑戰不在於模型能否獨立地辨識、回答、記憶或行動,而在於完整系統能否維持一個可靠、時間有效、可糾錯且可治理的感知-狀態-互動-行動迴圈。本綜述**首度以如此統一的框架系統性研究智慧眼鏡**。我們形式化第一人稱資料流與受限任務效用,沿八個可驗證的硬體能力軸刻畫裝置特性,圍繞七個相互關聯的基礎能力組織文獻,並引入涵蓋拍攝、反應式感知、情境輔助、持續狀態、受治理行動與具身耦合的L0-L5框架。在九個應用場景中,我們將任務與資料集、系統、產品、利害關係人、失效後果及證據缺口相連結。我們進一步提出九維度部署框架、以主張為條件的評估協議,以及從受控測量到縱向場域驗證與稽核的證據階梯。綜合而言,這些要素使智慧眼鏡更具可比較性、可部署性與可重現評估性,同時勾勒出邁向可信賴第一人稱智能的路線圖。
English
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop. This survey is the \textbf{first to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.