ChatPaper.aiChatPaper

EffectLearner: 実世界ビデオ物体除去のためのワールドアウェア物体効果推論

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

August 6, 2026
著者: Feier Wu, Wanke Xia, Xu He, Zilang Zhou, Si Chen, Dongxia Liu, Liyang Chen, Qimeng Wu, Zhengbo Zhang, Wenming Yang, Zhiyong Wu
cs.AI

要旨

ビデオ物体除去では、対象物体だけでなく、それによって誘発される影響も除去しつつ、高忠実度かつ時空間的に一貫した復元を維持する必要がある。既存手法は主に、事前定義された効果カテゴリと固定されたデータ分布から物体と効果の対応関係を暗黙的に学習するため、合成的な効果、空間的に分離または弱い相関しか持たない効果、ロングテールな物理現象、動的に進化する相互作用などを含む複雑な実世界シーンへの汎化が制限される。本論文では、意味推論を強化したフレームワークであるEffectLearnerを提案する。これは、VLMベースの物体・効果推論器とDiTベースのビデオ消去器を組み合わせたものである。構造化された効果分析プロンプトに導かれ、推論器は対象物体を強調したビデオに対してクロスモーダル推論を行い、コンパクトな効果認識コンテキストを抽出する。これにより、ビデオ消去器は包括的な物体・効果除去へと導かれる。さらに、動き認識マスクガイダンスと動き一貫性監視により、物体の動きやシーンの動的変化の下での除去範囲と時空間的安定性が向上する。挑戦的な実世界シナリオで本フレームワークを十分に活用するため、複雑な物体誘発効果に特化して設計されたペアビデオデータセットEffectWorldを構築し、一般的な監視と複雑効果データを組み合わせた段階的トレーニングカリキュラムを導入する。標準的なROSE-Benchでは、EffectLearnerは既存ベースラインをほとんどの指標で上回り、EffectWorld-Evalおよび挑戦的なEffectWorld-Wildの両方で明確な優位性を達成し、複雑な実世界シーンにおける高品質なビデオ物体除去の能力を示している。
English
Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.