見える嘘:身体的社会的相互作用におけるVLMエージェントによる言語・非言語を併用した欺瞞
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
August 31, 2026
著者: Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira, Gunhee Kim
cs.AI
要旨
LLMおよびVLMエージェントによる戦略的欺瞞は、AIアライメントと安全性における中核的懸念として浮上している。社会的推理ゲーム(各プレイヤーが秘匿された役割を保持し、他者とのコミュニケーションを通じて正体を推理するゲーム)は、特にマルチエージェント環境における標準的テストベッドとしての役割を担っている。しかしながら、既存のテストベッドはテキストのみで構成され、単一の固定エージェント構成で動作するため、欺瞞の分類体系が中核とみなす非言語的感覚運動チャネルが欠落しており、観察された行動が基盤となるモデルを反映しているのか、周辺のハーネスを反映しているのかが曖昧なままである。本研究では、インポスターエージェントが言語行動と非言語行動の統合を通じてクルーメイトを欺くことを要する3DマルチモーダルAmong UsサンドボックスであるMineAmongUsを導入する。さらに、5つの認知コンポーネントアブレーション軸を公開する設定可能なVLMエージェントハーネスであるARIAと、欺瞞の分類体系に基づき、ほぼ人間レベルのアトムラベル一致を達成するLLM-as-a-Judgeによって大規模に運用されるアトムレベルおよびアークレベルのアノテーションスキームを提案する。実証結果は、VLMエージェントが言語的欺瞞と非言語的欺瞞の統合を通じてインポスターの勝利を追求すること、そして非言語的チャネルがハーネスアブレーションとクロスVLM評価の両方にわたり、より決定的な勝利貢献要因として浮上することを示している。総括すると、本研究は身体化されたVLMエージェントのアライメント研究に新たな道を開くものである。
English
Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds, however, are text-only and run on a single fixed agent configuration, missing the non-verbal sensorimotor channels treated as core by deception taxonomies and leaving it ambiguous whether an observed behavior reflects the underlying model or the surrounding harness. We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action. We also propose ARIA, a configurable VLM-agent harness that exposes five cognitive-component ablation axes; and an atom- and arc-level annotation scheme grounded in deception taxonomies and operationalized at scale by an LLM-as-a-Judge reaching near-human atom-labeling agreement. Empirical results show that VLM agents pursue imposter wins through joint verbal and non-verbal deception, with non-verbal channels emerging as the more decisive winning contributors across both harness ablation and cross-VLM evaluation. Taken together, our work opens a new path for embodied VLM-agent alignment research.