Show-Harness: VLM 에이전트 하나만으로 로봇을 조종할 수 있다
Show-Harness: Just a VLM Agent Can Play Robots
September 9, 2026
저자: Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin, Yiqi Lin, Guoqiang Liang, Kevin Yuchen Ma, Qiming Huang, Mike Zheng Shou
cs.AI
초록
파운데이션 비전-언어 모델(VLM)은 세계에 대한 폭넓은 지능을 보이지만, 이러한 지능을 로봇 제어로 옮기는 일은 여전히 어렵다. 우리는 의도와 행동을 연결하는 간결한 의미 인터페이스를 통해 VLM이 로봇을 "플레이"할 수 있게 하는 체화형 하네스인 Show-Harness를 제안한다. Show-Harness는 VLM이 자연스럽게 추론할 수 있는 이산적 의미 행동 단위를 노출하는 한편, 임바디먼트별 해석기가 이를 결정론적으로 로컬 로봇 행동으로 그라운딩하도록 하여, VLM이 세밀한 물리적 결정을 직접 담당하도록 유지한다. 동일한 인터페이스를 통해 Show-Harness는 (1) 폐쇄형 프런티어 VLM을 제로샷 로봇 제어에 직접 활용하는 것과 (2) 소규모 오픈소스 VLM을 단 몇 GPU-시간의 미세조정만으로 저비용 배포에 적응시키는 것의 가능성을 입증한다. 우리는 또한 동일한 의미 행동 공간을 GUI 기반 시연 수집으로 확장하는 GUMI(GUI Manipulation Interface)를 개발한다. 이는 인간과 에이전트가 특수한 원격조작 하드웨어 없이 여러 임바디먼트에 걸쳐 로봇을 "플레이"할 수 있게 한다. 광범위한 실험은 Show-Harness를 장착한 VLM 에이전트가 과제, 임바디먼트, 환경 전반에 걸쳐 견고하게 일반화하며, 대표적인 에이전트형 및 VLA 패러다임을 능가함을 보여준다. 이러한 결과는 올바른 인터페이스가 추가 모델 용량이나 비용이 많이 드는 임바디먼트별 사전학습 없이도 파운데이션 VLM으로부터 상당한 체화 능력을 이끌어낼 수 있음을 시사한다.
English
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness demonstrates the feasibility of (1) directly unlocking closed-source frontier VLMs for zero-shot robot control, and (2) adapting small-scale open-source VLMs for low-cost deployment with just a few GPU-hours of fine-tuning. We further develop GUMI (GUI Manipulation Interface), which extends the same semantic action space to GUI-based demonstration collection, allowing humans and agents to "play" robots across embodiments without specialized teleoperation hardware. Extensive experiments show that Show-Harness-equipped VLM agents generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms. These results suggest that the right interface can unlock substantial embodied capability from foundation VLMs, without requiring additional model capacity or costly embodiment-specific pretraining.