세계로서의 코드: 물리적 추론을 위한 실행 가능한 세계 표현의 에이전트 기반 발견
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
August 27, 2026
저자: Hanyang Wang, Yimo Cai, Weiliang Chen, Jiawei Chi, Haowen Sun, Qiyu Dai, Yi-Hsin Hung, Xingzhuo Guo, Jinshan Ren, Runmao Yao, Ziwei Liu, Mingsheng Long, Yueqi Duan, Jun Gao, Jiangran Lyu, Fangfu Liu, Jialong Wu
cs.AI
초록
물리적 이해와 추론은 세상에 대한 간결하고 일반화 가능한 표현을 구축하는 데 의존한다. 현대의 비전-언어 모델은 다양한 물리적 사건을 인식하고 설명할 수 있지만, 세계가 어떻게 진화하고 개입에 어떻게 반응하는지를 안정적으로 추론하는 데 필요한 기저 메커니즘(객체 상태, 물리적 매개변수, 지배 역학 등)에 대한 명시적 표현이 부족한 경우가 많다. 본 연구에서는 실행 가능한 세계 표현을 통해 물리적 세계를 나타내는 패러다임인 Code-as-World를 소개한다. 물리적 구성, 동적 진화, 시각적 외형을 실행 가능한 코드로 표현함으로써, Code-as-World는 물리적 세계에 대한 간결하고 정량적 근거를 갖춘 제어 가능한 추상화를 제공한다. 자연어 설명이나 실세계 비디오와 같은 다중 모달 관측으로부터 이러한 표현을 구축하기 위해, 우리는 귀추적 추론에서 영감을 얻은 에이전트 기반 발견 루프를 개발한다. 이 루프에서 에이전트는 실행 가능한 세계 가설을 제안, 실행, 렌더링, 검증하고 반복적으로 정제한다. 구체적인 응용으로서, 우리는 검증된 실행 가능한 세계를 활용해 정량적 물리 추론을 위한 비전-언어 모델 훈련에 확장 가능한 물리적 감독을 제공한다. 실험 결과, Code-as-World-VL은 QuantiPhy에서 최고 수준의 성능을 달성하며 선도적인 독점 모델들을 능가하여, 물리적 지능을 위한 확장 가능한 기반으로서 실행 가능한 세계 표현의 잠재력을 보여준다.
English
Physical understanding and reasoning depend on forming compact and generalizable representations of the world. While modern vision-language models can recognize and explain diverse physical events, they often lack explicit representations of the underlying mechanisms-such as object states, physical parameters, and governing dynamics-needed for reliably reasoning how the world evolves and responds to interventions. In this work, we introduce Code-as-World, a paradigm that represents physical worlds through executable world representations. By expressing physical composition, dynamic evolution, and visual appearance as executable code, Code-as-World provides a compact, quantitatively grounded, and controllable abstraction of the physical world. To construct such representations from multimodal observations, such as natural-language descriptions or real-world videos, we develop an agentic discovery loop inspired by abductive reasoning, where an agent proposes, executes, renders, verifies, and iteratively refines executable world hypotheses. As a concrete application, we use verified executable worlds to provide scalable physical supervision for training vision-language models on quantitative physical reasoning. Experiments show that Code-as-World-VL achieves state-of-the-art performance on QuantiPhy and surpasses leading proprietary models, highlighting the potential of executable world representations as a scalable foundation for physical intelligence.