보기에서 행동으로: 스마트 글라스를 1인칭 지능 플랫폼으로
From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
August 25, 2026
저자: Jiangning Zhang, Haojun Chen, Yong Liu
cs.AI
초록
스마트 안경은 단순한 촬영 및 표시 액세서리에서 인간의 지각, 지속적 맥락, 디지털 및 물리적 행동을 연결하는 1인칭 지능 플랫폼으로 진화하고 있다. 이들의 신체 부착형 시점은 착용자의 시각, 청각, 운동, 손-객체 상호작용과 정렬되지만, 제한된 에너지, 열, 프라이버시, 피드백 제약 하에서 작동해야 한다. 증강 현실, 자아중심 시각, 멀티모달 모델, 인간-컴퓨터 상호작용, 체화 지능 분야에서 급속한 진전이 있었음에도 불구하고, 문헌은 장치, 작업, 벤치마크 전반에 걸쳐 분절된 상태로 남아 있다. 핵심 과제는 모델이 단독으로 인식, 응답, 기억, 행동할 수 있는지 여부가 아니라, 완전한 시스템이 신뢰할 수 있고 시간적으로 유효하며 교정 가능하고 관리 가능한 지각-상태-상호작용-행동 루프를 지속적으로 유지할 수 있는지 여부이다. 본 조사는 **이러한 통합 프레임워크를 통해 스마트 안경을 체계적으로 연구한 최초의 문헌**이다. 우리는 1인칭 데이터 흐름과 제약된 작업 효용성을 정형화하고, 8가지 검증 가능한 하드웨어 역량 축을 따라 장치를 특성화하며, 7가지 상호의존적 기반 역량을 중심으로 문헌을 체계화하고, 캡처, 반응형 지각, 맥락적 지원, 지속 상태, 관리된 행동, 체화적 결합을 포괄하는 L0-L5 프레임워크를 도입한다. 9가지 응용 시나리오에 걸쳐 작업을 데이터셋, 시스템, 제품, 이해관계자, 실패 결과, 증거 격차와 연결한다. 또한 9차원 배포 프레임워크, 주장 조건 기반 평가 프로토콜, 통제된 측정에서 종단적 현장 검증 및 감사에 이르는 증거 사다리를 제시한다. 이러한 요소들은 스마트 안경을 보다 비교 가능하고, 배포 가능하며, 재현 가능하게 평가할 수 있게 하는 동시에 신뢰할 수 있는 1인칭 지능을 향한 로드맵을 제시한다.
English
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop. This survey is the \textbf{first to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.