ChatPaper.aiChatPaper

Quo Vadis, 세계 모델링?

Quo Vadis, World Modeling?

August 3, 2026
저자: Yu Yang, Xuemeng Yang, Licheng Wen, Lingdong Kong, Xiaobin Hu, Dongyue Lu, Wei Chow, Xiyan Huang, Yuxiang Feng, Yue Liao, Jianbiao Mei, Daocheng Fu, Rong Wu, Pinlong Cai, Ran Yi, Ying Tai, Jiangning Zhang, Botian Shi, Yong Liu, Shuicheng Yan
cs.AI

초록

지속적으로 개선되는 에이전트는 정적 감독을 넘어 동적 상호작용 피드백을 필요로 하지만, 실제 환경과의 직접 상호작용은 비용이 많이 들고 느리며 안전하지 않고 병렬화하기 어렵다. 세계 모델링은 에이전트가 실제 행동을 취하기 전에 더 저렴하고 통제 가능한 피드백을 질의할 수 있게 해주는 자연스러운 중간 프록시를 제공한다. 고전적 세계 모델은 주로 미래 물리 상태 예측을 통해 이러한 프록시를 구현하는데, 이는 유용하지만 원시 상태 전이를 넘어 실행 가능한 피드백을 요구하는 에이전트에게는 좁은 형식화에 그친다. 본 연구에서는 에이전트 중심의 상호작용적 세계 프록시(Agent-Centric Interactive World Proxies)를 개념화하여, 기본 패러다임을 물리적 상태 전이에서 에이전트가 활용 가능한 정보 전이(예: 실행 결과, 검색된 경험 또는 기술, 검증 신호)로 전환함으로써 세계 모델링의 범위를 확장하고, 지속적으로 개선되는 에이전트에게 다용도 피드백을 제공하고자 한다. 이 설계 공간을 체계적으로 매핑하기 위해, 우리는 세계 프록시를 피드백 양식에 따라 여섯 가지 기능적 형태, 즉 동역학(dynamics), 공간(spatial), 실행(execution), 기억/경험(memory/experience), 기술(skill), 보상/검증(reward/verification) 프록시로 분류한다. 이는 세계 모델링이 에이전트 개선에 기여하는 주요 방식을 총체적으로 규명한다. 또한 우리는 이러한 프록시가 에이전트를 세 가지 점진적 수준에서 강화하는 방식을 분석한다. L.1(추론 시점 안내)에서는 프록시 출력이 맥락 내 정보를 풍부하게 하여 더 우수한 결정을 내리게 한다. L.2(훈련 시점 최적화)에서는 프록시 출력이 정책 학습을 위한 보상, 비평, 또는 합성 롤아웃을 생성한다. L.3(에이전트-프록시 공진화)에서는 실제 환경 증거가 프록시와 에이전트를 모두 지속적으로 업데이트하여 공진화를 이끈다. 궁극적으로 본 연구는 세계 모델링을 에이전트 중심 패러다임으로 재구성함으로써, 에이전트가 더 잘 계획하고, 더 빠르게 학습하며, 지속적으로 진화할 수 있게 하는 세계 프록시 구축을 위한 로드맵을 제시한다.
English
Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling offers a natural intermediate proxy that allows agents to query lower-cost, more controllable feedback before committing to real actions. Classical world models instantiate this proxy primarily through future physical-state prediction, a formulation useful yet narrow for agents that require actionable feedback beyond raw state transitions. In this work, we conceptualize Agent-Centric Interactive World Proxies, shifting the fundamental paradigm from physical state transitions to agent-usable information transitions, such as execution outcomes, retrieved experiences or skills, and verification signals, broadening the scope of world modeling to provide versatile feedback for continually improving agents. To systematically map this design space, we organize world proxies into six functional forms based on their feedback modalities: dynamics, spatial, execution, memory/experience, skill, and reward/verification proxies, which together characterize the primary ways world modeling serves agent improvement. We further analyze how these proxies empower agents across three progressive levels: L.1 Inference-Time Guidance, where proxy outputs enrich in-context information for superior decisions; L.2 Training-Time Optimization, where proxy outputs yield rewards, critiques, or synthetic rollouts for policy learning; and L.3 Agent-Proxy Co-Evolution, where real-environment evidence continuously updates both the proxy and the agent for co-evolution. Ultimately, this work recasts world modeling into an agent-centric paradigm, establishing a roadmap for building world proxies that empower agents to plan better, learn faster, and evolve continually.