VideoCoCo:以程式碼作為思維鏈,透過代理式雙引擎系統實現物理一致的影片生成
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
July 29, 2026
作者: Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Junyan Ye, Jinjing Zhao, Zehui Chen, Lin Chen, Renye Yan, Feng Zhao, Pheng-Ann Heng
cs.AI
摘要
文字轉影片模型已達到卓越的視覺品質,然而此類模型仍難以生成物理一致的動態,原因在於場景的時間演化必須從高度壓縮的文字提示中隱式推斷。現有的思維鏈方法引入了中間計畫或視覺狀態,但這些表徵通常不可執行或時間上過於稀疏,限制了其例示與控制完整時空過程的能力。為解決此限制,我們提出 VideoCoCo,一個代理式雙引擎框架,其中可執行的 Blender 程式碼作為過程層級的思維鏈。給定文字提示,編碼代理會合成一個 Blender 程式,明確指定場景及其時間演化。可執行的模擬引擎執行該程式以產生確定性的時空草稿,隨後由生成式影片引擎透過草稿條件化編輯將其轉化為照片級真實的影片。此分解方式將過程層級的推理與高保真度的視覺實現分離開來。為使影片編輯器適應模擬草稿,我們建構了 VideoCoCo-3K,一個經篩選的草稿-指令-目標三元組資料集。VideoCoCo 將 OmniWeaving 基線在 PhyGenBench 上從 0.475 提升至 0.558,在 VBench-2.0 上從 52.18 提升至 77.88,在兩個基準上均達到最佳平均分數。這些結果證明,可執行程式碼為物理一致的影片生成提供了有效、可控且可檢視的中間表徵。
English
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.