可见的谎言:VLM智能体在具身社交互动中的言语与非言语联合欺骗
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
August 31, 2026
作者: Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira, Gunhee Kim
cs.AI
摘要
大语言模型(LLM)与视觉语言模型(VLM)代理的策略性欺骗已成为AI对齐与安全领域的核心关注点。社交推理游戏(每位玩家持有隐藏身份,并通过相互交流来推断他人身份)构成了经典测试环境,尤其是在多智能体场景中。然而,现有测试环境仅支持文本模态,且运行于单一固定代理配置之上,既缺失了欺骗分类学所视为核心的非语言感觉运动通道,也使得观测到的行为究竟是反映底层模型还是外部框架变得模糊不清。我们提出了MineAmongUs——一个3D多模态《Among Us》沙盒环境,其中伪装者代理必须通过语言与非语言联合行动来欺骗船员。我们还提出了ARIA——一个可配置的VLM代理框架,暴露了五个认知组件消融轴;以及一个基于欺骗分类学、由LLM-as-a-Judge以接近人类水平的原子标签一致性进行规模化实施的原子级与弧级标注方案。实证结果表明,VLM代理通过语言与非语言联合欺骗来追求伪装者的胜利,且非语言通道在框架消融和跨VLM评估中均被证明是更具决定性的致胜因素。总体而言,我们的工作为具身VLM代理的对齐研究开辟了新路径。
English
Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds, however, are text-only and run on a single fixed agent configuration, missing the non-verbal sensorimotor channels treated as core by deception taxonomies and leaving it ambiguous whether an observed behavior reflects the underlying model or the surrounding harness. We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action. We also propose ARIA, a configurable VLM-agent harness that exposes five cognitive-component ablation axes; and an atom- and arc-level annotation scheme grounded in deception taxonomies and operationalized at scale by an LLM-as-a-Judge reaching near-human atom-labeling agreement. Empirical results show that VLM agents pursue imposter wins through joint verbal and non-verbal deception, with non-verbal channels emerging as the more decisive winning contributors across both harness ablation and cross-VLM evaluation. Taken together, our work opens a new path for embodied VLM-agent alignment research.