ChatPaper.aiChatPaper

Zetta ζ:自己進化する物理的知能のための効率的な閉ループ身体化ハーネス

Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence

August 17, 2026
著者: Xin Ding, Liang Mi, Mingzhe Huang, Zixuan Wang, Chao Zhang, Zixu Hao, Fu Chen, Xiangyu Li, Yikai Zheng, Yaoyu Guo, Weijun Wang, Kun Li, Hao Wu, Yunxin Liu, Ting Cao
cs.AI

要旨

身体化エージェントは、エンドツーエンドのポリシーモデルが残すギャップを埋めるためにますます利用されている。しかしながら、エージェント指向のアプローチは物理的実行における閉ループ学習を未だ実現していない。既存のハーネスは大部分がオープンループのままであり、ロールアウト中は固定されたスキルに従い、エピソード完了後にのみ内省を行う。このような事後的内省は、実行が展開する最中の制御を司ることはできない。なぜなら物理的相互作用は、今日の大規模エージェントモデルの処理頻度を超えるレートで急速に変化するロボットと環境の状態を追跡するための意思決定を必要とするからである。我々はZettaを提案する。Zettaは、ベースポリシーを凍結したまま、コードベースの実行時クリティックと回復スキルをオンラインで進化させる閉ループ身体化ハーネスである。時間スケールが分離された3つのループを通じて、Zettaはアクション頻度の統制、ロールアウトレベルのクリティック・回復スキル提案、および検証ゲート付きスキル更新を提供する。エージェントのロジックを異種の実行リソースから分離するロールアウト基盤であるZ-Infraと組み合わせることで、Zettaは現在のロールアウト予算の下でLIBERO-ProおよびRoboCasaにおいて最先端の成功率を達成し、それぞれ90.8%と93.6%に到達した。さらに、推論速度は11.1倍に高速化された。成功率は自己探索による経験とともにスケールし続け、学習されたスキルはゼロショットで転移し、明確なロボットの「アハ体験」が出現する。これらの結果は、閉ループハーネスの自己進化が、信頼性の高い物理的知能のためのスケーリング経路を切り拓くことを示している。
English
Embodied agents are increasingly used to close the gap left by end-to-end policy models. Yet the agentic path has not realized closed-loop learning in physical execution: existing harnesses remain largely open-loop, following fixed skills during rollout and reflecting only after an episode completes. Such post-hoc reflection cannot govern execution as it unfolds, because physical interaction requires decisions to track rapidly changing robot-environment states at a frequency beyond today's large agentic models. We present Zetta, a closed-loop embodied harness that evolves code-based runtime critics and recovery skills online while keeping the base policy frozen. Through three timescale-separated loops, Zetta provides action-frequency governance, rollout-level critic-recovery proposal, and validation-gated skill updates. Together with Z-Infra, a rollout infrastructure decoupling agent logic from heterogeneous execution resources, Zetta achieves state-of-the-art success on LIBERO-Pro and RoboCasa under our current rollout budget, reaching 90.8% and 93.6%, with an 11.1x inference speedup; success continues to scale with self-exploration experience; learned skills transfer zero-shot, and clear robotic "Aha Moments" emerge. These results show that closed-loop harness self-evolution opens a scaling path for reliable physical intelligence.