우리가 볼 수 있는 거짓말: 체화된 사회적 상호작용에서 VLM 에이전트의 언어적·비언어적 거짓말
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
August 31, 2026
저자: Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira, Gunhee Kim
cs.AI
초록
LLM 및 VLM 에이전트의 전략적 기만은 핵심적인 AI 정렬 및 안전 문제로 부상했다. 사회적 추론 게임(각 플레이어가 숨은 역할을 보유하고, 정체를 추론하기 위해 타인과 소통하는 게임)은 특히 다중 에이전트 환경에서 표준적인 시험대로 기능한다. 그러나 기존 테스트베드는 텍스트 전용이며 단일 고정 에이전트 구성에서 실행되어, 기만 분류 체계에서 핵심으로 간주되는 비언어적 감각운동 채널을 누락하고, 관찰된 행동이 기저 모델을 반영하는지 주변 하네스를 반영하는지 모호하게 만든다. 우리는 임포스터 에이전트가 언어적 행동과 비언어적 행동을 결합하여 크루메이트를 기만해야 하는 3D 다중모달 Among Us 샌드박스인 MineAmongUs를 소개한다. 또한 다섯 가지 인지 구성요소 절제 축을 노출하는 구성 가능한 VLM 에이전트 하네스인 ARIA와, 기만 분류 체계에 근거하고 LLM-as-a-Judge에 의해 대규모로 운영되어 인간 수준에 근접한 원자 레이블 일치도를 달성하는 원자 및 호 수준 주석 체계를 제안한다. 실증 결과는 VLM 에이전트가 언어적 및 비언어적 기만의 결합을 통해 임포스터 승리를 추구하며, 비언어적 채널이 하네스 절제와 교차 VLM 평가 모두에서 더 결정적인 승리 기여 요인으로 부상함을 보여준다. 종합적으로, 우리의 연구는 체화된 VLM 에이전트 정렬 연구를 위한 새로운 경로를 연다.
English
Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds, however, are text-only and run on a single fixed agent configuration, missing the non-verbal sensorimotor channels treated as core by deception taxonomies and leaving it ambiguous whether an observed behavior reflects the underlying model or the surrounding harness. We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action. We also propose ARIA, a configurable VLM-agent harness that exposes five cognitive-component ablation axes; and an atom- and arc-level annotation scheme grounded in deception taxonomies and operationalized at scale by an LLM-as-a-Judge reaching near-human atom-labeling agreement. Empirical results show that VLM agents pursue imposter wins through joint verbal and non-verbal deception, with non-verbal channels emerging as the more decisive winning contributors across both harness ablation and cross-VLM evaluation. Taken together, our work opens a new path for embodied VLM-agent alignment research.