ChatPaper.aiChatPaper

Zetta ζ:一種用於自演化物理智慧的高效閉環具身框架

Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence

August 17, 2026
作者: Xin Ding, Liang Mi, Mingzhe Huang, Zixuan Wang, Chao Zhang, Zixu Hao, Fu Chen, Xiangyu Li, Yikai Zheng, Yaoyu Guo, Weijun Wang, Kun Li, Hao Wu, Yunxin Liu, Ting Cao
cs.AI

摘要

具身智能體越來越常被用來填補端到端策略模型所留下的缺口。然而,智能體路徑尚未在物理執行中實現閉環學習:現有的框架在很大程度上仍屬開環,在執行期間遵循固定技能,僅在一個回合結束後進行反思。這種事後反思無法在執行展開的過程中進行調控,因為物理交互要求決策能夠以超越當今大型智能體模型的頻率,追蹤快速變化的機器人-環境狀態。我們提出Zetta,一個閉環具身框架,在保持基礎策略凍結的同時,線上演化基於程式碼的運行時評判器與復原技能。透過三個時間尺度分離的迴路,Zetta提供動作頻率調控、回合級評判-復原提案,以及驗證門控的技能更新。結合Z-Infra(一個將智能體邏輯與異構執行資源解耦的執行基礎設施),Zetta在我們當前的執行預算下於LIBERO-Pro和RoboCasa上達到了當前最優的成功率,分別達到90.8%和93.6%,並實現11.1倍的推理加速;成功率持續隨自我探索經驗而擴展;學習到的技能可零樣本遷移,並湧現出清晰的機器人「頓悟時刻」。這些結果表明,閉環框架的自我演化為可靠物理智能開闢了一條擴展路徑。
English
Embodied agents are increasingly used to close the gap left by end-to-end policy models. Yet the agentic path has not realized closed-loop learning in physical execution: existing harnesses remain largely open-loop, following fixed skills during rollout and reflecting only after an episode completes. Such post-hoc reflection cannot govern execution as it unfolds, because physical interaction requires decisions to track rapidly changing robot-environment states at a frequency beyond today's large agentic models. We present Zetta, a closed-loop embodied harness that evolves code-based runtime critics and recovery skills online while keeping the base policy frozen. Through three timescale-separated loops, Zetta provides action-frequency governance, rollout-level critic-recovery proposal, and validation-gated skill updates. Together with Z-Infra, a rollout infrastructure decoupling agent logic from heterogeneous execution resources, Zetta achieves state-of-the-art success on LIBERO-Pro and RoboCasa under our current rollout budget, reaching 90.8% and 93.6%, with an 11.1x inference speedup; success continues to scale with self-exploration experience; learned skills transfer zero-shot, and clear robotic "Aha Moments" emerge. These results show that closed-loop harness self-evolution opens a scaling path for reliable physical intelligence.