Hallo4D: マルチモーダル幻覚軽減による一貫時空間生成
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
July 15, 2026
著者: Hongbo Wang, Huaibo Huang, Jie Cao, Jin Liu, Haoyang Tong, Ran He
cs.AI
要旨
近年の3D生成技術の進歩により、視覚的合成において顕著な成果が得られているが、既存手法の多くは幾何学的整合性を明示的に保証する機構を持たない2D拡散の教師信号に依存しており、構造の重複や位置ずれなどの空間的幻覚を引き起こす。こうした問題は4D生成においてさらに深刻化し、視点間および時間的変化にわたる整合性維持に伴う課題として、ジッター、アイデンティティのフリッカー、構造的ドリフトなどが生じる。本稿では、3Dおよび4Dコンテンツ生成における時空間的幻覚を軽減するための統一的かつモデル非依存のフレームワーク「Hallo4D」を提案する。Hallo4Dは、生成・検出・修正のパラダイムを導入し、大規模マルチモーダル言語モデル(LMM)を活用して、マルチビューおよびマルチフレームレンダリングから空間的・時間的な不一致を特定・要約する。これらの知見に基づき、LMMベースのセレクタがマルチモデル投票を通じて修正候補を評価するコンセンサス駆動型の画像空間整合性最適化を導く。本手法は再学習やアーキテクチャ変更を必要としない。さらに、時間的整合性と最適化効率を向上させるため、Hallo4Dは動作認識キーフレームサンプリング、LMM誘導型初期化、外観アライメントを組み込む。また、露出認識最適化と可視性プルーニングを導入し、困難な視点下でのロバスト性を強化する。広範な実験により、Hallo4Dは多様な3D・4D生成設定において強力なベースラインを一貫して上回り、整合性を考慮したコンテンツ生成に対するスケーラブルで汎用的なソリューションを提供することを実証する。
English
While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated structures and misaligned geometry. These issues become more severe in 4D generation, where maintaining consistency across viewpoints and temporal evolution introduces additional challenges, including jitter, identity flicker, and structural drift. We present Hallo4D, a unified and model-agnostic framework for mitigating spatiotemporal hallucinations in 3D and 4D content generation. Hallo4D introduces a generation-detection-correction paradigm that leverages large multimodal language models (LMMs) to identify and summarize spatial and temporal inconsistencies from multi-view and multi-frame renderings. These insights guide a consensus-driven image-space consistency optimization, where an LMM-based selector evaluates candidate corrections through multi-model voting, without requiring retraining or architectural modifications. To further improve temporal consistency and optimization efficiency, Hallo4D incorporates motion-aware keyframe sampling, LMM-guided initialization, and appearance alignment. We additionally introduce exposure-aware optimization and visibility pruning to enhance robustness under challenging viewpoints. Extensive experiments demonstrate that Hallo4D consistently outperforms strong baselines across diverse 3D and 4D generation settings, providing a scalable and generalizable solution for consistency-aware content generation.