ChatPaper.aiChatPaper

Hallo4D: Multimodale hallucinatiemitigatie voor consistente spatio-temporele generatie

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

July 15, 2026
Auteurs: Hongbo Wang, Huaibo Huang, Jie Cao, Jin Liu, Haoyang Tong, Ran He
cs.AI

Samenvatting

Hoewel recente vooruitgang in 3D-generatie indrukwekkende visuele synthese mogelijk heeft gemaakt, vertrouwen bestaande methoden vaak op 2D-diffusiesupervisie zonder expliciete mechanismen voor geometrische consistentie, wat leidt tot ruimtelijke hallucinaties zoals gedupliceerde structuren en niet-uitgelijnde geometrie. Deze problemen worden ernstiger bij 4D-generatie, waar het handhaven van consistentie over gezichtspunten en temporele evolutie extra uitdagingen met zich meebrengt, zoals jitter, identiteitsflikkering en structurele drift. We presenteren Hallo4D, een uniform en model-agnostisch raamwerk voor het verminderen van ruimtelijk-temporele hallucinaties bij 3D- en 4D-inhoudsgeneratie. Hallo4D introduceert een generatie-detectie-correctieparadigma dat gebruikmaakt van grote multimodale taalmodellen (LMM's) om ruimtelijke en temporele inconsistenties uit multi-view- en multi-frame-renderingen te identificeren en samen te vatten. Deze inzichten sturen een consensusgedreven optimalisatie van consistentie in de beeldruimte, waarbij een op LMM gebaseerde selector kandidaatcorrecties evalueert via multi-modelstemming, zonder dat hertraining of architecturale aanpassingen nodig zijn. Om de temporele consistentie en optimalisatie-efficiëntie verder te verbeteren, bevat Hallo4D bewegingsbewuste keyframe-sampling, LMM-gestuurde initialisatie en uiterlijkuitlijning. We introduceren daarnaast belichtingsbewuste optimalisatie en zichtbaarheidspruning om de robuustheid onder uitdagende gezichtspunten te vergroten. Uitgebreide experimenten tonen aan dat Hallo4D consequent beter presteert dan sterke baselines in diverse 3D- en 4D-generatie-instellingen, en een schaalbare en generaliseerbare oplossing biedt voor consistentiebewuste inhoudsgeneratie.
English
While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated structures and misaligned geometry. These issues become more severe in 4D generation, where maintaining consistency across viewpoints and temporal evolution introduces additional challenges, including jitter, identity flicker, and structural drift. We present Hallo4D, a unified and model-agnostic framework for mitigating spatiotemporal hallucinations in 3D and 4D content generation. Hallo4D introduces a generation-detection-correction paradigm that leverages large multimodal language models (LMMs) to identify and summarize spatial and temporal inconsistencies from multi-view and multi-frame renderings. These insights guide a consensus-driven image-space consistency optimization, where an LMM-based selector evaluates candidate corrections through multi-model voting, without requiring retraining or architectural modifications. To further improve temporal consistency and optimization efficiency, Hallo4D incorporates motion-aware keyframe sampling, LMM-guided initialization, and appearance alignment. We additionally introduce exposure-aware optimization and visibility pruning to enhance robustness under challenging viewpoints. Extensive experiments demonstrate that Hallo4D consistently outperforms strong baselines across diverse 3D and 4D generation settings, providing a scalable and generalizable solution for consistency-aware content generation.