ChatPaper.aiChatPaper

从感知到行动:智能眼镜作为第一人称智能平台

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms

August 25, 2026
作者: Jiangning Zhang, Haojun Chen, Yong Liu
cs.AI

摘要

智能眼镜正从采集与显示配件演进为第一人称智能平台,将人类感知、持续性情境以及数字或物理行动连接起来。其穿戴式视角与佩戴者的视觉、听觉、运动及手-物交互保持一致,但必须在严格的能量、热功耗、隐私和反馈约束下运行。尽管增强现实、自我中心视觉、多模态模型、人机交互和具身智能取得了快速发展,相关文献在设备、任务和基准测试方面仍呈现碎片化状态。关键挑战不在于模型能否在孤立条件下完成识别、回答、记忆或行动,而在于一个完整系统能否维持可靠、时间有效、可纠正且可治理的感知-状态-交互-行动闭环。本综述是**首个通过此类统一框架系统研究智能眼镜的工作**。我们形式化定义了第一人称数据流和受限任务效用,沿八条可验证的硬件能力轴刻画设备特征,围绕七项相互依赖的基础能力组织文献体系,并引入涵盖采集、反应式感知、情境辅助、持续性状态、受治理行动和具身耦合的L0-L5框架。在九个应用场景中,我们将任务与数据集、系统、产品、利益相关者、失效后果和证据缺口相连接。我们进一步提出了九维部署框架、基于主张条件的评估协议,以及从受控测量到纵向现场验证和审计的证据阶梯。这些要素共同使智能眼镜更具可比性、可部署性和可复现评估性,同时勾勒出通往可信第一人称智能的路线图。
English
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop. This survey is the \textbf{first to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.