ChatPaper.aiChatPaper

持続的ストーリーとインタラクティブな世界のための長期的音声・映像生成

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

August 24, 2026
著者: Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang
cs.AI

要旨

ビデオ生成は、孤立したクリップから長編ナラティブやインタラクティブなワールドへと進化しており、モデルにはアイデンティティの保持、ユーザー制御への追従、長時間のロールアウトにわたる安定性が求められている。本稿では、2つの目的別バリアントを備えた統合音声-視覚生成システムJoyAI-Echo-1.5を提案する。長編ビデオバリアントは、複数の先行ショットにわたる視覚的証拠を集約するコンポーザブルなクロスショットメモリと、音声フィルタリングを施したショット全体の音声から導出される話者キューを導入し、テキスト、画像、メモリ条件付けの柔軟な組み合わせにわたって、持続的なキャラクター外観と音声アイデンティティを可能にする。ワールドモデルバリアントは、異種のナビゲーション入力を較正されたメトリック6自由度カメラ軌道に変換し、幾何学認識条件付け経路を通じて注入することで、柔軟な視点にわたるコントローラ非依存のインタラクションを実現する。長時間ホライズン生成を効率的にサポートするため、双方向音声-視覚バックボーンを、プログレッシブ教師フォーシングと、自己生成ロールアウトに対する短期・長期ホライズンのSelf-Gradient Forcingを用いて、因果的少数ステップ生成器へと変換する。実験により、両設定において優れた性能が実証された。JoyAI-Echo-1.5は、既存の長編ビデオベースラインと比較して、クロスショット一貫性、視覚品質、テキスト整合性、音声忠実度の各指標で改善を達成している。ワールドモデルバリアントは、平均スコア81.7でWBenchで第1位を獲得し、SANA-WM-Benchで最高の視覚品質と長期ホライズン持続性を達成している。これらの結果は、メモリ、幾何学的制御、ロールアウトを考慮した学習が、一貫性のあるストーリーと継続的に進化するインタラクティブなワールドを生成するための実用的な基盤を提供することを示している。プロジェクトページ: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/
English
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.