ChatPaper.aiChatPaper

我們能看見的謊言:視覺語言模型代理在具身社交互動中的聯合言語與非言語欺騙

Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions

August 31, 2026
作者: Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira, Gunhee Kim
cs.AI

摘要

LLM與VLM智能體的策略性欺騙行為已成為AI對齊與安全領域的核心關注點。社交推理遊戲(每位玩家持有隱藏角色,並透過相互溝通推斷身分)長期作為此類研究的標準測試平台,尤其在多智能體環境中。然而,現有測試平台僅限於純文字形式,且運行於單一固定的智能體配置之上,既遺漏了欺騙分類法視為核心的非語言感覺運動通道,亦無法釐清所觀察到的行為究竟反映底層模型本身抑或其所處的測試框架。我們提出MineAmongUs——一個3D多模態《Among Us》沙盒環境,其中內鬼智能體必須透過語言與非語言行動的協同配合來欺騙船員。此外,我們提出ARIA——一個可配置的VLM智能體測試框架,揭示五個認知元件的消融維度;並提出一套植基於欺騙分類法的原子級與弧級標註方案,透過以LLM作為裁判(LLM-as-a-Judge)的方式大規模落地,其原子標註一致性接近人類水準。實證結果顯示,VLM智能體確實會透過語言與非語言欺騙的協同配合來追求內鬼勝利,其中非語言通道在框架消融與跨VLM評估中均浮現為更具決定性的致勝因素。綜上所述,本研究為具身VLM智能體對齊研究開闢了新路徑。
English
Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds, however, are text-only and run on a single fixed agent configuration, missing the non-verbal sensorimotor channels treated as core by deception taxonomies and leaving it ambiguous whether an observed behavior reflects the underlying model or the surrounding harness. We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action. We also propose ARIA, a configurable VLM-agent harness that exposes five cognitive-component ablation axes; and an atom- and arc-level annotation scheme grounded in deception taxonomies and operationalized at scale by an LLM-as-a-Judge reaching near-human atom-labeling agreement. Empirical results show that VLM agents pursue imposter wins through joint verbal and non-verbal deception, with non-verbal channels emerging as the more decisive winning contributors across both harness ablation and cross-VLM evaluation. Taken together, our work opens a new path for embodied VLM-agent alignment research.