ChatPaper.aiChatPaper

HumanCLAW: 視覚言語モデルは身体を通して行動できるか?

HumanCLAW: Can Vision-Language Models Act Through a Body?

July 29, 2026
著者: Siyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li, Ziwei Liu, Chuan Guo
cs.AI

要旨

視覚言語モデル(VLM)が物理的な身体を通じて行動できるかを評価することは困難である。行動の結果は、VLMの判断とモーター制御が結合したものとなる。タスクが失敗した場合、VLMが誤った選択をしたのか、単にモーター制御がそれを実行できなかったのか(例えば、バランスを崩して転倒した場合など)を区別するのは難しい。本研究では、行動決定を低レベルの実行から分離する評価フレームワークHumanCLAWを提案する。各ステップにおいて、装着された既製のVLMがアトミックスキルコマンドを発行し、そのコマンドは重力や衝突を含む現実の物理的影響を伴う、秒未満の連続的な全身動作に変換される。これにより、身体は物理世界で自由に行動できる一方で、バランスやモーターエラーといった実行側の外乱は除外される。測定可能となるのは、モデルの行動知能、すなわち身体が次に何を実行すべきかという瞬間ごとの選択である。このフレームワークに基づき、HumanCLAW-Benchを構築した。これは、41の屋内シーンにわたる1,218件の長期的かつ一人称視点の「発見・移動・操作」エピソードからなる。9つの最先端VLMをテストした結果、いずれもベンチマークを解くことはできず、最良のモデルでも成功率はわずか16.8%であった。目標物体の認識がボトルネックではない。現在のVLMに欠けているのは身体化された自己認識であり、自身の身体がどこにあるのか、目標に到達したかどうか、障害物に衝突したかどうかを把握できず、見失ってしまうのである。
English
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.