ChatPaper.aiChatPaper

HumanCLAW: 비전-언어 모델은 몸을 통해 행동할 수 있는가?

HumanCLAW: Can Vision-Language Models Act Through a Body?

July 29, 2026
저자: Siyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li, Ziwei Liu, Chuan Guo
cs.AI

초록

시각-언어 모델(VLM)이 물리적 신체를 통해 행동할 수 있는지 평가하는 것은 까다로운 과제이다. 행동의 결과는 VLM의 결정과 모터 제어가 결합된 것이다. 작업이 실패했을 때, VLM이 잘못된 선택을 했는지, 아니면 모터 컨트롤러가 단순히 이를 실행하지 못했는지(예: 균형을 잃고 넘어짐) 알기 어렵다. 본 연구에서는 행동 결정을 저수준 실행과 분리하는 평가 프레임워크인 HumanCLAW를 소개한다. 각 단계마다, 하네스에 고정된 기성 VLM이 원자적 기술 명령을 발행하면, 해당 명령은 중력 및 충돌을 포함한 실제 물리적 결과를 수반하는 1초 미만 단위의 연속적인 전신 동작으로 변환된다. 따라서 신체는 물리적 세계에서 자유롭게 행동할 수 있으면서도, 실행 측의 교란, 균형 및 모터 오류는 배제된다. 측정 가능한 것은 모델의 행동 지능, 즉 매 순간 신체가 다음에 무엇을 실행해야 하는지에 대한 선택이다. 이 프레임워크를 기반으로 HumanCLAW-Bench를 구축했다: 41개의 실내 장면에 걸친 1,218개의 장기적이고 1인칭 시점의 찾기-이동-상호작용 에피소드이다. 최신 VLM 9개를 테스트한 결과, 어떤 모델도 벤치마크를 해결하지 못했으며, 최고 성능 모델조차 성공률 16.8%에 그쳤다. 대상을 인식하는 것은 병목이 아니다. 현재 VLM이 부족한 것은 체화된 자기 인식이다. 즉, 자신의 신체 위치, 목표에 도달했는지 여부, 장애물에 부딪혔는지 여부를 추적하지 못한다.
English
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.