SceneActBench: エージェントは見た3Dシーンに対して行動できるか?
SceneActBench: Can Agents Act on the 3D Scenes They See?
July 24, 2026
著者: Yifei Zhao, Xiangxin Zhou, Wenhao Yang, Jiaqi Tang, Pu Jian, Huanjin Yao, Jiarui Yao, Haowei Lin, Chunchao Guo, Zhuo Chen, Wenkai Lyu, Jianzhu Ma, Xueqian Wang, Wenxi Zhu
cs.AI
要旨
視覚言語モデル(VLM)エージェントは、3Dシーンを記述するだけでなく、それらに対して行動するためにツールを利用するようになってきている。既存の3Dベンチマークはテキスト応答や単一オブジェクト操作を評価しており、複数オブジェクトからなる完全な3Dシーンにおけるエージェントの行動は評価が不十分なままである。我々はSceneActBenchを提案する。これは、統一されたエージェント・環境ループのもとでの5つの3Dタスクにわたる、視覚に基づく行動のベンチマークである。PNG画像またはサンプリングされたビデオフレーム、および該当する場合は提供された3Dアセットが与えられると、エージェントは3D環境で行動する。各最終出力を、タスク固有の幾何学的指標を用いて隠された正解データと照合して評価する。SceneActBenchは、210のソースインスタンスから構築された5つのタスクから成り、ペアとなる入力条件を含む520のタスクケースを生成する。比較の公平性を保つため、すべてのタスクは1つの固定されたエージェントループを通じて実行される。11のプロプライエタリなVLM構成において、総合スコアは38.6から50.2の範囲にわたり、どの構成もタスク間で一貫して良好なパフォーマンスを示さなかった。我々はさらに、失敗がどこでどのように発生するかを分析する。
English
Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where applicable, supplied 3D assets, an agent acts on a 3D environment. We evaluate each final output against hidden ground truth with task-specific geometric metrics. SceneActBench comprises five tasks built from 210 source instances, yielding 520 task cases including paired input conditions. Every task runs through one fixed agent loop to keep the comparison fair. Across eleven proprietary VLM configurations, Overall scores span 38.6-50.2, and none performs consistently well across tasks. We further analyse where and how failures manifest.