ChatPaper.aiChatPaper

Marigold V2: 단안 깊이 추정을 위한 확산 트랜스포머 재조명

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

September 8, 2026
저자: Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine Süsstrunk, Dengxin Dai
cs.AI

초록

단안 깊이 추정은 어디에나 적용되는 동시에 고도로 ill-posed한 컴퓨터 비전 과제로, 장면 재구성, 계산 사진학, 로보틱스 등에 하위 응용된다. 이 분야가 성숙했음에도 불구하고, 최근 모델들은 여전히 분포 외(out-of-distribution) 입력에 일반화하고 선명하고 세밀한 깊이 맵을 생성하는 데 어려움을 겪는다. 본 논문에서 우리는 확산 트랜스포머(DiT) 아키텍처로 구동되는 현대 이미지 생성 및 편집 모델을 최첨단 단안 깊이 추정기로 재활용하기 위한 일련의 기법인 Marigold를 다시 살펴본다. 우리의 레시피는 사전 학습된 다단계 흐름 매칭(flow-matching) 모델로부터의 단일 단계 추론을 목표로 하며, 필요할 경우 양자화를 적용하여 모델 용량을 보존하면서도 실행 비용을 낮게 유지한다. 우리는 단순 학습(naive training)의 아티팩트를 분석하고 두 가지 효과적인 해결책을 확인한다: 모델의 내부 표현을 정답(ground-truth)에서 추출한 의미 특징과 정렬하는 것, 그리고 새로운 Sinkhorn 기반 손실을 중심으로 구성된 2단계 미세 조정 프로토콜을 채택하는 것이다. 그 결과, KITTI 및 ETH3D에서 이전 최고 성능 대비 AbsRel이 16-26% 향상되면서 분포 외에서도 잘 일반화되는 더 선명하고 깔끔한 깊이 맵이 얻어진다. 정성적으로, 우리 모델은 기존 모델들이 포착하지 못했던 털, 잎사귀, 머리카락처럼 얇은 경계를 해결한다. 나아가, Marigold V2는 표면 법선 추정 및 고유 이미지 분해와 같은 다른 조밀 회귀 과제에 적용될 때 최첨단 결과를 달성한다. 프로젝트 웹사이트: https://hf.co/spaces/huawei-bayerlab/marigold-v2-web
English
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp and detailed depth maps. In this paper, we revisit Marigold, a set of techniques for repurposing modern image generation and editing models, powered by the diffusion transformer (DiT) architecture, into state-of-the-art monocular depth estimators. Our recipes target single-step inference from pretrained multi-step flow-matching models, with quantization where needed, preserving model capacity while remaining cheap to run. We analyze the artifacts of naive training and identify two effective remedies: aligning the model's internal representations with semantic features extracted from ground-truth, and adopting a 2-stage fine-tuning protocol built around a novel Sinkhorn-based loss. The results are crisper, cleaner depth maps that generalize well out-of-distribution, with 16-26% improvement in AbsRel over the previous best on KITTI and ETH3D. Qualitatively, our model resolves fur, foliage, and hair-thin edges that have eluded prior models. Furthermore, Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition. Project website: https://hf.co/spaces/huawei-bayerlab/marigold-v2-web