HumanCLAW:視覺語言模型能否透過身體行動?
HumanCLAW: Can Vision-Language Models Act Through a Body?
July 29, 2026
作者: Siyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li, Ziwei Liu, Chuan Guo
cs.AI
摘要
評估一個視覺-語言模型是否能透過身體來行動是一項挑戰。行動的結果結合了視覺-語言模型的決策與運動控制。當任務失敗時,難以判斷是視覺-語言模型做出了錯誤選擇,還是運動控制器根本無法執行(例如失去平衡而跌倒)。在這項工作中,我們提出 HumanCLAW,這是一個能將行動決策與低階執行解耦的評估框架。在每一步中,一個被固定好的現成視覺-語言模型會發出一個原子技能指令,而該指令會被轉化為一段次秒級的連續全身動作,並伴隨真實物理效果(包括重力與碰撞)。因此,身體可以在物理世界中自由行動,同時執行層面的干擾(如平衡與運動誤差)則被排除在外。最終可測量的是模型的行動智能:它在每一刻選擇身體下一步該執行什麼的能力。基於此框架,我們建立了 HumanCLAW-Bench:涵蓋 41 個室內場景中 1,218 個長時域、以自我為中心的「尋找-導航-互動」片段。我們測試了九個最新的視覺-語言模型,發現沒有任何一個能通過這項基準測試;表現最佳的模型僅達到 16.8% 的成功率。辨識目標並非瓶頸。當前視覺-語言模型所欠缺的是具身自我意識:它們會遺忘自身身體的位置,無法判斷自己在哪裡、是否已抵達目標、或者是否已撞上障礙物。
English
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.