ChatPaper.aiChatPaper

OracleZoom:受同策略自蒸餾啟發之參考約束遞迴影像超解析度

OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution

September 6, 2026
作者: Shubhashis Roy Dipta, Sourajit Saha, Shaswati Saha, Nobin Sarwar
cs.AI

摘要

遞迴超解析度(SR)透過將預測反覆饋回同一模型,將固定尺度 SR 擴展至極端放大,類似於反覆放大一張影像。然而,要在每個尺度皆取得真值,尤其在深層,仍具挑戰性,因為所需的來源解析度呈幾何級數成長,使更深層的預測缺乏監督。我們提出 OracleZoom,一個受同策略蒸餾啟發、受參考約束的框架,能在其軌跡上訓練,同時將最後的真值證據帶到監督邊界之外。直接與跨尺度監督約束可驗證內容,而無參考品質目標則引導尚未解析的細尺度細節。受 KL 約束的預訓練潛在先驗限制由品質驅動的漂移,而 EMA 一致性則穩定監督邊界。在七個資料集上,OracleZoom 在跨縮放尺度上達到最先進的 SR 品質,平均為 0.713 CLIPIQA,且在更深尺度上有更大增益,同時顯著減少幻覺。程式碼、資料與模型可於 https://dipta007.github.io/OracleZoom/ 取得。
English
Recursive Super-Resolution (SR) extends fixed-scale SR to extreme magnification by repeatedly feeding predictions back into the same model, analogous to zooming an image repeatedly. However, ground truth availability at every scale, especially at depth, remains challenging as the required source resolution grows geometrically, leaving deeper predictions unsupervised. We present OracleZoom, an on-policy distillation-inspired, reference-constrained framework that trains on its trajectory while carrying the last ground-truth evidence beyond the supervision boundary. Direct and cross-scale supervision constrain verifiable content, while a no-reference quality objective guides unresolved fine-scale detail. A KL-constrained pretrained latent prior limits quality-driven drift, while EMA consistency stabilizes the supervision boundary. Across seven datasets, OracleZoom achieves the state-of-the-art SR quality across zooming scales, averaging 0.713 CLIPIQA, with larger gains on deeper scales, while significantly reducing hallucinations. Code, data, and models are available at https://dipta007.github.io/OracleZoom/ .