ChatPaper.aiChatPaper

VideoCoCo: エージェント型デュアルエンジンシステムによる物理的に一貫した映像生成のためのコードとしてのCoT

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

July 29, 2026
著者: Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Junyan Ye, Jinjing Zhao, Zehui Chen, Lin Chen, Renye Yan, Feng Zhao, Pheng-Ann Heng
cs.AI

要旨

テキストから動画を生成するモデルは目覚ましい視覚的品質を達成しているが、シーンの時間的発展が高度に圧縮されたテキストプロンプトから暗黙的に推論されなければならないため、物理的に一貫したダイナミクスを生成することには依然として課題が残る。既存のチェーン・オブ・ソート(思考連鎖)手法は中間的な計画や視覚状態を導入するが、これらの表現は通常、実行不可能であるか時間的に疎であり、完全な時空間プロセスを具体化し制御する能力を制限している。この限界に対処するため、我々はVideoCoCoを導入する。これはエージェント型のデュアルエンジンフレームワークであり、実行可能なBlenderコードがプロセスレベルの思考連鎖として機能する。テキストプロンプトが与えられると、コーディングエージェントがシーンとその時間的発展を明示的に指定するBlenderプログラムを合成する。実行可能なシミュレーションエンジンがプログラムを実行して決定的な時空間ドラフトを生成し、その後、生成型ビデオエンジンがドラフトに条件付けられた編集によってこれをフォトリアリスティックな動画に変換する。この分解により、プロセスレベルの推論と高忠実度の視覚的実現が分離される。ビデオエディタをシミュレーションによるドラフトに適応させるために、我々はドラフト・指示・ターゲットのトリプレットからなる厳選データセットであるVideoCoCo-3Kを構築した。VideoCoCoは、PhyGenBenchにおいてOmniWeavingベースラインを0.475から0.558へ、VBench-2.0において52.18から77.88へ改善し、両ベンチマークで最高平均スコアを達成した。これらの結果は、実行可能なコードが物理的に一貫した動画生成のための、効果的で制御可能かつ検査可能な中間表現を提供することを示している。
English
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.