ChatPaper.aiChatPaper

Zetta ζ:面向自演化物理智能的高效闭环具身框架

Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence

August 17, 2026
作者: Xin Ding, Liang Mi, Mingzhe Huang, Zixuan Wang, Chao Zhang, Zixu Hao, Fu Chen, Xiangyu Li, Yikai Zheng, Yaoyu Guo, Weijun Wang, Kun Li, Hao Wu, Yunxin Liu, Ting Cao
cs.AI

摘要

具身智能体越来越多地被用于弥合端到端策略模型留下的差距。然而,智能体路径尚未在物理执行中实现闭环学习:现有框架在很大程度上仍是开环的,在rollout期间遵循固定技能,并且仅在一个回合结束后才进行反思。这种事后反思无法在执行过程中对其进行治理,因为物理交互要求决策以超越当今大型智能体模型的频率来跟踪快速变化的机器人-环境状态。我们提出了Zetta,一种闭环具身框架,它在保持基础策略冻结的同时,在线演化基于代码的运行时评论器和恢复技能。通过三个时间尺度分离的环路,Zetta提供了动作频率治理、rollout级评论器-恢复提案以及验证门控的技能更新。与Z-Infra(一种将智能体逻辑与异构执行资源解耦的rollout基础设施)一起,Zetta在当前rollout预算下在LIBERO-Pro和RoboCasa上取得了最先进的成功,分别达到90.8%和93.6%,推理速度提升11.1倍;成功随自我探索经验持续扩展;学习到的技能可零样本迁移,并出现了清晰的机器人“顿悟时刻”。这些结果表明,闭环框架的自我演化为可靠的物理智能开辟了一条扩展路径。
English
Embodied agents are increasingly used to close the gap left by end-to-end policy models. Yet the agentic path has not realized closed-loop learning in physical execution: existing harnesses remain largely open-loop, following fixed skills during rollout and reflecting only after an episode completes. Such post-hoc reflection cannot govern execution as it unfolds, because physical interaction requires decisions to track rapidly changing robot-environment states at a frequency beyond today's large agentic models. We present Zetta, a closed-loop embodied harness that evolves code-based runtime critics and recovery skills online while keeping the base policy frozen. Through three timescale-separated loops, Zetta provides action-frequency governance, rollout-level critic-recovery proposal, and validation-gated skill updates. Together with Z-Infra, a rollout infrastructure decoupling agent logic from heterogeneous execution resources, Zetta achieves state-of-the-art success on LIBERO-Pro and RoboCasa under our current rollout budget, reaching 90.8% and 93.6%, with an 11.1x inference speedup; success continues to scale with self-exploration experience; learned skills transfer zero-shot, and clear robotic "Aha Moments" emerge. These results show that closed-loop harness self-evolution opens a scaling path for reliable physical intelligence.