ChatPaper.aiChatPaper

Hallo4D:多模態幻覺緩解以實現一致時空生成

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

July 15, 2026
作者: Hongbo Wang, Huaibo Huang, Jie Cao, Jin Liu, Haoyang Tong, Ran He
cs.AI

摘要

近年來,三維生成技術的進展已能實現令人驚豔的視覺合成,然而現有方法多依賴二維擴散監督,缺乏明確的幾何一致性機制,因而容易產生空間幻覺,例如重複結構與幾何錯位。這些問題在四維生成中更加嚴重,因為跨視角與時間演進的一致性維護帶來了額外挑戰,包括抖動、身份閃爍與結構漂移。我們提出 Hallo4D,一個統一且與模型無關的框架,用於減緩三維與四維內容生成中的時空幻覺。Hallo4D 導入「生成-檢測-校正」範式,利用大型多模態語言模型(LMM)從多視角與多幀渲染中識別並總結空間與時間上的不一致性。這些洞察驅動一個共識導向的影像空間一致性最佳化,其中基於 LMM 的選擇器透過多模型投票評估候選校正,無需重新訓練或修改架構。為進一步提升時間一致性與最佳化效率,Hallo4D 整合了動作感知關鍵幀取樣、LMM 引導初始化以及外觀對齊。此外,我們引入曝光感知最佳化與可見性剪枝,以增強在具挑戰性視角下的穩健性。大量實驗顯示,Hallo4D 在各種三維與四維生成設定中一致優於強基線,為一致性導向的內容生成提供了一個可擴展且具泛化能力的解決方案。
English
While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated structures and misaligned geometry. These issues become more severe in 4D generation, where maintaining consistency across viewpoints and temporal evolution introduces additional challenges, including jitter, identity flicker, and structural drift. We present Hallo4D, a unified and model-agnostic framework for mitigating spatiotemporal hallucinations in 3D and 4D content generation. Hallo4D introduces a generation-detection-correction paradigm that leverages large multimodal language models (LMMs) to identify and summarize spatial and temporal inconsistencies from multi-view and multi-frame renderings. These insights guide a consensus-driven image-space consistency optimization, where an LMM-based selector evaluates candidate corrections through multi-model voting, without requiring retraining or architectural modifications. To further improve temporal consistency and optimization efficiency, Hallo4D incorporates motion-aware keyframe sampling, LMM-guided initialization, and appearance alignment. We additionally introduce exposure-aware optimization and visibility pruning to enhance robustness under challenging viewpoints. Extensive experiments demonstrate that Hallo4D consistently outperforms strong baselines across diverse 3D and 4D generation settings, providing a scalable and generalizable solution for consistency-aware content generation.