ReFlowSET: SAR-EO 이미지 변환을 위한 표현 정렬 잠재 흐름 매칭
ReFlowSET: Representation-Aligned Latent Flow Matching for SAR-to-EO Image Translation
September 1, 2026
저자: Jeonghyeok Do, Seungchul Lee, Munchurl Kim
cs.AI
초록
SAR-to-EO 이미지 변환은 합성개구레이더(SAR) 관측으로부터 전자광학(EO) 영상을 생성하는 것을 목표로 한다. 기존의 잠재 확산 접근법들은 일반적으로 사전에 정의된 오토인코더를 그대로 계승해 사용하지만, 재구성 충실도는 코덱과 모달리티에 따라 크게 달라질 수 있다. 잠재 코덱은 SAR 조건과 EO 대상 모두의 왕복 보존에 영향을 미치기 때문에 코덱 선택은 근본적인 설계 결정에 해당한다. 그럼에도 기존 방법들은 대부분 자연 이미지에 대해 사전 학습된 코덱에 주로 의존해 왔다. 이러한 문제를 해결하기 위해, 우리는 SAR-EO 공동 재구성 검증을 통해 코덱을 선택하는 조건부 잠재 플로우 매칭 프레임워크인 ReFlowSET을 제안한다. ReFlowSET은 무거운 사전 학습 생성기를 계승하는 대신, 선택된 잠재 공간에서 현저히 더 작은 조건부 DiT를 처음부터 학습하며, 이중 스트림 SAR 조건화와 후속 공동 특징 정제를 사용한다. 이러한 처음부터의 학습에 의미론적 안내를 제공하기 위해, 중간 단계의 노이즈가 포함된 EO 특징을 동결된 비전 파운데이션 모델에서 추출한 잡음 없는 대상 EO 표현과 정렬한다. 이 정렬은 학습 중에만 사용되며 추론 시 추가 비용을 발생시키지 않는다. QXS-SAROPT 및 SAR2Opt 데이터셋 실험은 다양한 지각 충실도 및 분포 지표에서 최첨단 성능을 보여준다. 코드와 사전 학습 가중치는 https://github.com/KAIST-VICLab/ReFlowSET에서 공개적으로 제공된다.
English
SAR-to-EO image translation aims to generate electro-optical (EO) imagery from synthetic aperture radar (SAR) observations. Existing latent diffusion approaches typically inherit a predetermined autoencoder, although reconstruction fidelity can vary substantially across codecs and modalities. Because the latent codec affects the round-trip preservation of both SAR conditions and EO targets, codec selection constitutes a fundamental design choice; nevertheless, existing methods largely rely on codecs pretrained on natural images. To remedy this, we introduce ReFlowSET, a conditional latent flow-matching framework that selects its codec through a joint SAR--EO reconstruction audit. Rather than inheriting a heavyweight pretrained generator, ReFlowSET trains a substantially smaller conditional DiT from scratch in the selected latent space, using dual-stream SAR conditioning followed by joint feature refinement. To provide semantic guidance for this from-scratch training, intermediate noisy-EO features are aligned with clean target-EO representations extracted by a frozen vision foundation model. This alignment is used only during training and introduces no additional inference cost. Experiments on QXS-SAROPT and SAR2Opt demonstrate state-of-the-art performance across diverse perceptual fidelity and distributional metrics. Code and pretrained weights are publicly available at https://github.com/KAIST-VICLab/ReFlowSET.