長時程音訊-視覺生成:用於持續性故事與互動式世界
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
August 24, 2026
作者: Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang
cs.AI
摘要
影片生成正從孤立的片段進展到長篇敘事與互動式世界,這要求模型能夠保持角色身分、遵循使用者控制,並在長時間推演中維持穩定性。我們提出 JoyAI-Echo-1.5,一個統一的音訊-視覺生成系統,包含兩個專用變體。長影片變體引入了可組合的跨鏡頭記憶,該記憶會彙整多個先前鏡頭的視覺證據,以及從語音過濾後的全鏡頭音訊中提取的語者線索,從而在文字、影像與記憶條件的靈活組合下,實現持久的角色外觀與語音身分。世界模型變體將異質導航輸入轉換為校準的公制六自由度相機軌跡,並透過幾何感知的條件控制路徑注入,從而在靈活視角下實現與控制器無關的互動。為了支援高效能的長時程生成,我們利用漸進式教師強制,以及在自生成展開序列上的短時程與長時程自梯度強制,將雙向音訊-視覺骨幹網路轉換為因果少步生成器。實驗證明在兩種設定下均展現出強大的效能。JoyAI-Echo-1.5 在跨鏡頭一致性、視覺品質、文字對齊與語音保真度方面均優於現有的長影片基線模型。其世界模型變體在 WBench 上排名第一,平均得分為 81.7,並在 SANA-WM-Bench 上取得領先的視覺品質與長時程持久性。綜合來看,這些結果表明記憶、幾何控制與展開感知訓練為生成連貫的故事與持續演化的互動世界提供了實用的基礎。專案頁面:https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/。
English
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.