Marigold V2: 単眼深度推定のための拡散トランスフォーマーの再検討
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
September 8, 2026
著者: Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine Süsstrunk, Dengxin Dai
cs.AI
要旨
単眼深度推定は、広く研究されている一方で、本質的に高度な不良設定問題であるコンピュータビジョンタスクであり、シーン再構成、コンピュテーショナルフォトグラフィ、ロボティクスなどに下流応用を持つ。この分野は成熟しているにもかかわらず、最近のモデルでさえ、分布外入力への汎化や、鮮明で詳細な深度マップの生成には依然として苦戦している。本論文では、拡散トランスフォーマ(DiT)アーキテクチャを基盤とする現代の画像生成・編集モデルを、最先端の単眼深度推定器へと転用するための一連の技術であるMarigoldを再検討する。我々の手法は、必要に応じた量子化を伴いつつ、事前学習済みの多段階フローマッチングモデルからの1ステップ推論を目指し、モデル容量を維持しながら実行コストを低く抑える。我々は、単純な学習に伴うアーティファクトを分析し、2つの有効な対策を特定する。すなわち、モデルの内部表現を正解データから抽出した意味特徴量と整合させることと、新規のSinkhornベース損失を中心に構築した2段階ファインチューニングプロトコルを採用することである。その結果として、より鮮明でクリーンな深度マップが得られ、分布外に対して良好に汎化し、KITTIおよびETH3Dにおいて以前の最高性能に対してAbsRelを16〜26%改善する。定性的には、我々のモデルは、従来モデルでは捉えられなかった毛皮、葉群、髪の毛のように細いエッジを再現する。さらに、Marigold V2は、表面法線推定や本質画像分解など、他の密回帰タスクに適用した場合にも最先端の結果を達成する。プロジェクトウェブサイト:https://hf.co/spaces/huawei-bayerlab/marigold-v2-web
English
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp and detailed depth maps. In this paper, we revisit Marigold, a set of techniques for repurposing modern image generation and editing models, powered by the diffusion transformer (DiT) architecture, into state-of-the-art monocular depth estimators. Our recipes target single-step inference from pretrained multi-step flow-matching models, with quantization where needed, preserving model capacity while remaining cheap to run. We analyze the artifacts of naive training and identify two effective remedies: aligning the model's internal representations with semantic features extracted from ground-truth, and adopting a 2-stage fine-tuning protocol built around a novel Sinkhorn-based loss. The results are crisper, cleaner depth maps that generalize well out-of-distribution, with 16-26% improvement in AbsRel over the previous best on KITTI and ETH3D. Qualitatively, our model resolves fur, foliage, and hair-thin edges that have eluded prior models. Furthermore, Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition. Project website: https://hf.co/spaces/huawei-bayerlab/marigold-v2-web