見ることから行動へ:一人称知能プラットフォームとしてのスマートグラス
From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
August 25, 2026
著者: Jiangning Zhang, Haojun Chen, Yong Liu
cs.AI
要旨
スマートグラスは、単なる撮影・表示アクセサリから、人間の知覚、持続的な文脈、デジタルまたは物理的な行動を接続する一人称知能プラットフォームへと進化しつつある。装着型視点は着用者の視覚、聴覚、運動、手と物体のインタラクションに一致する一方で、厳しいエネルギー、熱、プライバシー、フィードバックの制約下で動作する必要がある。拡張現実、一人称視覚、マルチモーダルモデル、人間とコンピュータのインタラクション、身体化知能における急速な進歩にもかかわらず、既存の研究はデバイス、タスク、ベンチマーク間で断片化したままである。核心的な課題は、モデルが単独で認識、回答、記憶、行動できるかどうかではなく、完全なシステムが信頼性が高く、時間的に有効で、訂正可能かつ統治可能な知覚-状態-相互作用-行動ループを維持できるかどうかにある。本サーベイは、**このような統一フレームワークを通じてスマートグラスを系統的に研究する初めてのもの**である。我々は一人称データフローと制約付きタスク実用性を形式化し、検証可能な8つのハードウェア機能軸に沿ってデバイスを特徴付け、7つの相互依存する基盤能力に基づいて文献を整理し、キャプチャ、反応的知覚、文脈支援、持続状態、統治された行動、身体化結合を網羅するL0-L5フレームワークを導入する。9つの応用シーンにわたり、タスクをデータセット、システム、製品、利害関係者、障害の結果、エビデンスのギャップと結び付ける。さらに、9次元の展開フレームワーク、主張条件付き評価プロトコル、および管理された測定から縦断的フィールド検証と監査に至るエビデンスの階層を提示する。これらの要素を組み合わせることで、スマートグラスの比較可能性、展開可能性、再現可能な評価可能性が向上し、信頼できる一人称知能へのロードマップが示される。
English
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop. This survey is the \textbf{first to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.