ChatPaper.aiChatPaper

Show-Harness:僅憑一個 VLM 代理即可操控機器人

Show-Harness: Just a VLM Agent Can Play Robots

September 9, 2026
作者: Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin, Yiqi Lin, Guoqiang Liang, Kevin Yuchen Ma, Qiming Huang, Mike Zheng Shou
cs.AI

摘要

基礎視覺語言模型(VLM)展現出對世界的廣泛智能,然而要將這種智能轉化為機器人控制仍具挑戰性。我們提出 Show-Harness,一套具身化框架,能讓 VLM 透過一個緊湊的語意介面「玩」機器人,將意圖連結至行動。Show-Harness 提供離散的語意動作單元,讓 VLM 能自然進行推理;而具身專屬解譯器則以確定性方式將這些單元接地為該具身上的機器人動作,使 VLM 直接負責細粒度的物理決策。透過相同介面,Show-Harness 展示了以下可行性:(1) 直接解鎖閉源前沿 VLM 以進行零樣本機器人控制;(2) 僅需數個 GPU 小時的微調,即可調適小規模開源 VLM 以實現低成本部署。我們進一步開發 GUMI(GUI 操作介面),將相同的語意動作空間延伸至基於 GUI 的示範收集,讓人類與代理能在無需專門遙操作硬體的情況下,跨具身「玩」機器人。大量實驗顯示,配備 Show-Harness 的 VLM 代理能在任務、具身與環境之間穩健泛化,表現優於代表性的代理式與 VLA 範式。這些結果表明,正確的介面能從基礎 VLM 中釋放可觀的具身能力,而無需額外的模型容量或昂貴的具身專屬預訓練。
English
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness demonstrates the feasibility of (1) directly unlocking closed-source frontier VLMs for zero-shot robot control, and (2) adapting small-scale open-source VLMs for low-cost deployment with just a few GPU-hours of fine-tuning. We further develop GUMI (GUI Manipulation Interface), which extends the same semantic action space to GUI-based demonstration collection, allowing humans and agents to "play" robots across embodiments without specialized teleoperation hardware. Extensive experiments show that Show-Harness-equipped VLM agents generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms. These results suggest that the right interface can unlock substantial embodied capability from foundation VLMs, without requiring additional model capacity or costly embodiment-specific pretraining.