VideoCoCo: 에이전트 기반 이중 엔진 시스템을 통한 물리적으로 일관된 비디오 생성을 위한 Code-as-CoT
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
July 29, 2026
저자: Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Junyan Ye, Jinjing Zhao, Zehui Chen, Lin Chen, Renye Yan, Feng Zhao, Pheng-Ann Heng
cs.AI
초록
텍스트-비디오 모델은 놀라운 시각적 품질을 달성했지만, 장면의 시간적 전개가 고도로 압축된 텍스트 프롬프트로부터 암시적으로 추론되어야 하기 때문에 물리적으로 일관된 역학을 생성하는 데 여전히 어려움을 겪는다. 기존의 사고 사슬(chain-of-thought) 접근법은 중간 계획이나 시각적 상태를 도입하지만, 이러한 표현은 일반적으로 실행 가능하지 않거나 시간적으로 희소하여 완전한 시공간 과정을 구체화하고 제어하는 능력을 제한한다. 이 한계를 해결하기 위해, 우리는 실행 가능한 Blender 코드가 프로세스 수준의 사고 사슬 역할을 하는 에이전트 기반 이중 엔진 프레임워크인 VideoCoCo를 제안한다. 텍스트 프롬프트가 주어지면, 코딩 에이전트가 장면과 그 시간적 전개를 명시적으로 지정하는 Blender 프로그램을 합성한다. 실행 가능한 시뮬레이션 엔진은 프로그램을 실행하여 결정론적 시공간 초안을 생성하고, 이후 생성 비디오 엔진이 초안 조건화 편집(draft-conditioned editing)을 통해 이를 포토리얼리스틱 비디오로 변환한다. 이러한 분해는 프로세스 수준 추론과 고충실도 시각적 구현을 분리한다. 시뮬레이션된 초안에 비디오 편집기를 적응시키기 위해, 우리는 초안-지시-대상(draft-instruction-target) 삼중 항목으로 구성된 큐레이션 데이터셋인 VideoCoCo-3K를 구축한다. VideoCoCo는 PhyGenBench에서 OmniWeaving 기준선을 0.475에서 0.558로, VBench-2.0에서는 52.18에서 77.88로 향상시켜 두 벤치마크 모두에서 최고 평균 점수를 달성한다. 이러한 결과는 실행 가능한 코드가 물리적으로 일관된 비디오 생성을 위한 효과적이고 제어 가능하며 검사 가능한 중간 표현을 제공함을 보여준다.
English
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.