지속적인 스토리와 인터랙티브 월드를 위한 장기적 시청각 생성
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
August 24, 2026
저자: Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang
cs.AI
초록
비디오 생성은 개별 클립을 넘어 장편 서사와 인터랙티브 월드로 발전하고 있으며, 모델이 정체성을 보존하고 사용자 제어를 따르며 확장된 롤아웃에서 안정성을 유지해야 한다. 우리는 두 가지 특수 목적 변형을 갖춘 통합 오디오-비주얼 생성 시스템인 JoyAI-Echo-1.5를 제시한다. 장편 비디오 변형은 여러 이전 샷에 걸쳐 시각적 증거를 집계하고 음성 필터링된 전체 샷 오디오에서 추출된 화자 신호를 활용하는 구성 가능한 크로스샷 메모리를 도입하여, 텍스트, 이미지, 메모리 컨디셔닝의 유연한 조합에서 지속적인 캐릭터 외형과 음성 정체성을 가능하게 한다. 월드 모델 변형은 이기종 내비게이션 입력을 보정된 미터법 6-DoF 카메라 궤적으로 변환하고 이를 기하학 인지 컨디셔닝 경로를 통해 주입함으로써, 유연한 시점에서 컨트롤러 무관 상호작용을 가능하게 한다. 효율적인 장기 생성을 지원하기 위해, 우리는 점진적 티처 포싱과 자기 생성 롤아웃에 대한 단기 및 장기 Self-Gradient Forcing을 사용하여 양방향 오디오-비주얼 백본을 인과적 소수 단계 생성기로 변환한다. 실험 결과는 두 설정 모두에서 강력한 성능을 입증한다. JoyAI-Echo-1.5는 기존 장편 비디오 기준선 대비 샷 간 일관성, 시각적 품질, 텍스트 정렬, 음성 충실도에서 개선을 달성한다. 월드 모델 변형은 평균 점수 81.7로 WBench에서 1위를 차지하며, SANA-WM-Bench에서 최고 수준의 시각적 품질과 장기 지속성을 달성한다. 종합하면, 이러한 결과는 메모리, 기하학적 제어, 롤아웃 인지 학습이 일관된 이야기와 지속적으로 진화하는 인터랙티브 월드를 생성하는 데 실용적 기반을 제공함을 시사한다. 프로젝트 페이지: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.
English
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.