ChatPaper.aiChatPaper

SceneActBench:智能体能否对其所见的三维场景进行操作?

SceneActBench: Can Agents Act on the 3D Scenes They See?

July 24, 2026
作者: Yifei Zhao, Xiangxin Zhou, Wenhao Yang, Jiaqi Tang, Pu Jian, Huanjin Yao, Jiarui Yao, Haowei Lin, Chunchao Guo, Zhuo Chen, Wenkai Lyu, Jianzhu Ma, Xueqian Wang, Wenxi Zhu
cs.AI

摘要

视觉语言模型(VLM)智能体越来越多地借助工具在3D场景中执行操作,而不仅限于描述场景。现有3D基准测试仅对文本响应或单物体操作进行评分,导致智能体在完整多物体3D场景中的操作能力缺乏评估。我们提出SceneActBench基准测试,在统一智能体-环境循环框架下对五种3D任务的视觉条件化操作能力进行评估。智能体基于PNG图像或采样视频帧(如适用,还可结合提供的3D资产)对3D环境执行操作。我们采用任务特定的几何度量指标,将每个最终输出与隐藏真实值进行对比评估。SceneActBench包含210个源实例构建的五种任务,形成包含配对输入条件的520个任务案例。为确保公平比较,每个任务均通过单一固定智能体循环执行。在十一种专有VLM配置中,总体得分范围为38.6-50.2,且没有任何配置能在所有任务中持续表现优异。我们进一步分析了故障发生的具体情境与成因。
English
Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where applicable, supplied 3D assets, an agent acts on a 3D environment. We evaluate each final output against hidden ground truth with task-specific geometric metrics. SceneActBench comprises five tasks built from 210 source instances, yielding 520 task cases including paired input conditions. Every task runs through one fixed agent loop to keep the comparison fair. Across eleven proprietary VLM configurations, Overall scores span 38.6-50.2, and none performs consistently well across tasks. We further analyse where and how failures manifest.