ChatPaper.aiChatPaper

Zetta ζ: 자가 진화 물리적 지능을 위한 효율적인 폐루프 체화 프레임워크

Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence

August 17, 2026
저자: Xin Ding, Liang Mi, Mingzhe Huang, Zixuan Wang, Chao Zhang, Zixu Hao, Fu Chen, Xiangyu Li, Yikai Zheng, Yaoyu Guo, Weijun Wang, Kun Li, Hao Wu, Yunxin Liu, Ting Cao
cs.AI

초록

체화된 에이전트(embodied agent)는 종단간(end-to-end) 정책 모델이 남긴 한계를 메우기 위해 점점 더 많이 활용되고 있다. 그러나 에이전트적 접근은 물리적 실행에서 폐루프 학습(closed-loop learning)을 실현하지 못했다. 기존의 하네스(harness)는 대체로 개루프(open-loop)로 유지되어, 롤아웃 중에는 고정된 스킬을 따르고 에피소드가 끝난 후에만 반성한다. 이러한 사후 반성은 실행이 진행되는 동안 이를 제어할 수 없는데, 물리적 상호작용은 오늘날의 대형 에이전트 모델이 감당할 수 있는 빈도를 넘어서는 속도로 변화하는 로봇-환경 상태를 추적하는 결정을 요구하기 때문이다. 우리는 기본 정책을 동결시킨 채 코드 기반 런타임 비평가(critic)와 복구 스킬을 온라인으로 진화시키는 폐루프 체화 하네스인 Zetta를 제시한다. 세 가지 시간 척도로 분리된 루프를 통해 Zetta는 행동 빈도 제어, 롤아웃 수준의 비평가-복구 제안, 검증 게이트를 적용한 스킬 업데이트를 제공한다. Zetta는 에이전트 로직을 이기종 실행 자원으로부터 분리하는 롤아웃 인프라인 Z-Infra와 함께, 현재 롤아웃 예산 하에서 LIBERO-Pro와 RoboCasa에서 최첨단 성공률을 달성하여 각각 90.8%와 93.6%에 도달하고 11.1배의 추론 속도 향상을 보인다. 성공률은 자기 탐험 경험에 따라 계속 확장되며, 학습된 스킬은 제로샷으로 전이되고, 명확한 로봇의 ‘아하 순간(Aha Moment)’이 나타난다. 이러한 결과는 폐루프 하네스의 자기 진화가 신뢰할 수 있는 물리적 지능을 위한 확장 경로를 연다는 것을 보여준다.
English
Embodied agents are increasingly used to close the gap left by end-to-end policy models. Yet the agentic path has not realized closed-loop learning in physical execution: existing harnesses remain largely open-loop, following fixed skills during rollout and reflecting only after an episode completes. Such post-hoc reflection cannot govern execution as it unfolds, because physical interaction requires decisions to track rapidly changing robot-environment states at a frequency beyond today's large agentic models. We present Zetta, a closed-loop embodied harness that evolves code-based runtime critics and recovery skills online while keeping the base policy frozen. Through three timescale-separated loops, Zetta provides action-frequency governance, rollout-level critic-recovery proposal, and validation-gated skill updates. Together with Z-Infra, a rollout infrastructure decoupling agent logic from heterogeneous execution resources, Zetta achieves state-of-the-art success on LIBERO-Pro and RoboCasa under our current rollout budget, reaching 90.8% and 93.6%, with an 11.1x inference speedup; success continues to scale with self-exploration experience; learned skills transfer zero-shot, and clear robotic "Aha Moments" emerge. These results show that closed-loop harness self-evolution opens a scaling path for reliable physical intelligence.