에이전트 기반 게임 개발: 세계 모델 확장을 위한 검증 가능한 궤적 데이터 엔진

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

August 26, 2026
저자: Pengfei Zhou, Hexin Wang, Zhengfeiyang Zhang, Yixing Ma, Zhenglin Wan, Kaipeng Zhang, Wangbo Zhao, Yang You
cs.AI

초록

세계 모델을 확장하는 일반적인 전략은 더 많은 연산을 투입해 더 많은 크롤링 비디오로 훈련하는 것이다. 우리는 이러한 전략이 비효율적이라고 주장한다. 세계 모델의 확장에는 근거 있는 보상 신호를 제공하는 재귀적 데이터 엔진 또한 필요하기 때문이다. 코드 에이전트의 성공은 이것이 왜 중요한지를 잘 보여준다. 코드는 실행 가능하므로, 컴파일러와 런타임이 LLM의 강화학습(RL) 사후 훈련에 고품질의 보상을 제공할 수 있다. 반면 공간 생성은 여전히 CLIP 점수와 같은 모호한 프록시에 크게 의존한다. 이러한 신호는 불명확하고 편향되어 있어 RL 사후 훈련을 지원하기 어렵다. 이에 비해 게임 개발은 공간 세계 모델에 대해 누락된 보상 환경을 제공한다. 게임 엔진에 의해 인코딩된 장면은 실행 가능한 세계 명세이다. 엔진은 충돌, 물리, 이동 가능성, 제한된 플레이 가능성을 효율적으로 검사할 수 있으며, 개발자는 장면의 수용 여부를 판단하여 전역 검증 신호를 제공한다. 또한 게임 개발은 RL 사후 훈련을 위한 현실적인 장기 궤적 데이터를 제공한다. 이에 우리는 밀집된 엔진 신호와 개발 과정에서 발생하는 암묵적 인간 수용 피드백을 결합한 사후 훈련 패러다임인 인간-엔진 검증을 통한 강화학습(RLHEV)을 제안한다.
English
A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers grounded reward signals. The success of code agents illustrates why this matters. As code is executable, compilers and runtimes can provide high-quality rewards for Reinforcement Learning (RL) post-training of LLMs. By contrast, spatial generation still relies largely on fuzzy proxies such as CLIP scores. These signals are fuzzy and biased, making them hard to support RL post-training. Compared with these, game development provides a missing reward environment for spatial world models. A scene encoded by a game engine is an executable world specification: the engine can efficiently check collision, physics, navigability and bounded playability, while the developer provides the global verification signal by judging whether the scene should be accepted. Game development also provides real-world long-horizon trajectory data for RL post-training. We therefore propose Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process.
PDF1181August 29, 2026