ChatPaper.aiChatPaper

面向持久故事与交互世界的长时程视听生成

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

August 24, 2026
作者: Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang
cs.AI

摘要

视频生成正超越孤立片段,向长篇叙事和交互式世界发展,这要求模型保持角色身份一致性、遵循用户控制,并在长时间生成过程中保持稳定。我们提出了JoyAI-Echo-1.5,一个统一的音视频生成系统,包含两个专门设计的变体。长视频变体引入了可组合的跨镜头记忆机制,能够聚合多个先前镜头中的视觉证据,以及从语音过滤的全镜头音频中提取的说话人线索,从而在文本、图像和记忆条件的灵活组合下实现角色外观和语音身份的持续保持。世界模型变体将异构的导航输入转换为校准的度量式六自由度相机轨迹,并通过几何感知条件通路进行注入,实现跨灵活视角的控制器无关交互。为支持高效的长时程生成,我们将双向音视频骨干网络转变为因果式少步生成器,采用渐进式教师强制以及基于自生成展开的短期和长期自梯度强制。实验表明,该系统在两种设置下均表现优异。JoyAI-Echo-1.5在跨镜头一致性、视觉质量、文本对齐和语音保真度方面均优于现有长视频基线。其世界模型变体在WBench上排名第一,平均得分81.7,并在SANA-WM-Bench上取得领先的视觉质量和长时间持续性。综合而言,这些结果表明,记忆机制、几何控制与展开感知训练为生成连贯故事和持续演化的交互式世界提供了实用基础。项目主页:https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/。
English
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.