Hallo4D: 일관된 시공간 생성을 위한 다중 모드 환각 완화
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
July 15, 2026
저자: Hongbo Wang, Huaibo Huang, Jie Cao, Jin Liu, Haoyang Tong, Ran He
cs.AI
초록
최근 3D 생성 분야의 발전으로 인상적인 시각적 합성이 가능해졌지만, 기존 방법들은 종종 기하학적 일관성을 위한 명시적 메커니즘 없이 2D 확산 감독에 의존하여 중복 구조나 정렬되지 않은 기하학과 같은 공간적 환각을 초래한다. 이러한 문제는 4D 생성에서 더욱 심각해지는데, 시점 간 및 시간적 진화에 따른 일관성 유지가 떨림, 정체성 깜빡임, 구조적 표류와 같은 추가적 과제를 도입하기 때문이다. 본 논문에서는 3D 및 4D 콘텐츠 생성에서 시공간적 환각을 완화하기 위한 통합적이고 모델에 구애받지 않는 프레임워크인 Hallo4D를 제안한다. Hallo4D는 생성-탐지-교정 패러다임을 도입하여, 대규모 다중 모달 언어 모델(LMM)을 활용해 다중 시점 및 다중 프레임 렌더링에서 공간적 및 시간적 불일치를 식별하고 요약한다. 이러한 통찰은 합의 기반 이미지 공간 일관성 최적화를 안내하며, 여기서 LMM 기반 선택기가 다중 모델 투표를 통해 후보 교정을 평가하되 재학습이나 아키텍처 수정을 요구하지 않는다. 시간적 일관성과 최적화 효율성을 더욱 개선하기 위해 Hallo4D는 움직임 인식 키프레임 샘플링, LMM 안내 초기화, 외형 정렬을 통합한다. 또한 까다로운 시점에서의 강건성을 강화하기 위해 노출 인지 최적화와 가시성 가지치기를 추가로 도입한다. 광범위한 실험을 통해 Hallo4D가 다양한 3D 및 4D 생성 설정에서 강력한 기준 모델들을 일관되게 능가하며, 일관성 인식 콘텐츠 생성을 위한 확장 가능하고 일반화 가능한 솔루션을 제공함을 입증한다.
English
While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated structures and misaligned geometry. These issues become more severe in 4D generation, where maintaining consistency across viewpoints and temporal evolution introduces additional challenges, including jitter, identity flicker, and structural drift. We present Hallo4D, a unified and model-agnostic framework for mitigating spatiotemporal hallucinations in 3D and 4D content generation. Hallo4D introduces a generation-detection-correction paradigm that leverages large multimodal language models (LMMs) to identify and summarize spatial and temporal inconsistencies from multi-view and multi-frame renderings. These insights guide a consensus-driven image-space consistency optimization, where an LMM-based selector evaluates candidate corrections through multi-model voting, without requiring retraining or architectural modifications. To further improve temporal consistency and optimization efficiency, Hallo4D incorporates motion-aware keyframe sampling, LMM-guided initialization, and appearance alignment. We additionally introduce exposure-aware optimization and visibility pruning to enhance robustness under challenging viewpoints. Extensive experiments demonstrate that Hallo4D consistently outperforms strong baselines across diverse 3D and 4D generation settings, providing a scalable and generalizable solution for consistency-aware content generation.