OracleZoom: 온폴리시 자기 증류에서 착안한 참조 제약 재귀적 이미지 초해상도
OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution
September 6, 2026
저자: Shubhashis Roy Dipta, Sourajit Saha, Shaswati Saha, Nobin Sarwar
cs.AI
초록
재귀적 초해상도(Recursive Super-Resolution, SR)는 예측을 동일 모델에 반복적으로 피드백하여 고정 스케일 SR을 극한 확대까지 확장하며, 이는 이미지를 반복적으로 확대하는 것과 유사하다. 그러나 모든 스케일에서, 특히 깊은 스케일에서 정답(ground truth)의 가용성은 요구되는 소스 해상도가 기하급수적으로 증가하기 때문에 여전히 어려운 과제로 남아 있으며, 더 깊은 예측은 비지도 상태로 남게 된다. 우리는 OracleZoom을 제안한다. 이는 온-폴리시 증류에서 영감을 받은 참조 제약 프레임워크로, 자신의 궤적에서 학습하면서 마지막 정답 증거를 지도 경계 너머까지 전달한다. 직접 및 교차 스케일 지도는 검증 가능한 콘텐츠를 제약하고, 무참조 품질 목적 함수는 해결되지 않은 미세 스케일 디테일을 안내한다. KL 제약된 사전 학습 잠재 사전은 품질 주도 드리프트를 제한하며, EMA 일관성은 지도 경계를 안정화한다. 7개 데이터셋에서 OracleZoom은 줌 스케일 전반에 걸쳐 최고 수준의 SR 품질을 달성하여 평균 0.713 CLIPIQA를 기록하고, 더 깊은 스케일에서 더 큰 향상을 보이며, 환각을 상당히 줄인다. 코드, 데이터, 모델은 https://dipta007.github.io/OracleZoom/ 에서 이용할 수 있다.
English
Recursive Super-Resolution (SR) extends fixed-scale SR to extreme magnification by repeatedly feeding predictions back into the same model, analogous to zooming an image repeatedly. However, ground truth availability at every scale, especially at depth, remains challenging as the required source resolution grows geometrically, leaving deeper predictions unsupervised. We present OracleZoom, an on-policy distillation-inspired, reference-constrained framework that trains on its trajectory while carrying the last ground-truth evidence beyond the supervision boundary. Direct and cross-scale supervision constrain verifiable content, while a no-reference quality objective guides unresolved fine-scale detail. A KL-constrained pretrained latent prior limits quality-driven drift, while EMA consistency stabilizes the supervision boundary. Across seven datasets, OracleZoom achieves the state-of-the-art SR quality across zooming scales, averaging 0.713 CLIPIQA, with larger gains on deeper scales, while significantly reducing hallucinations. Code, data, and models are available at https://dipta007.github.io/OracleZoom/ .