ChatPaper.aiChatPaper

EDITBRIDGE: 충실하고 효율적인 초고해상도 이미지 편집을 위한

EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

August 18, 2026
저자: Jiayi Song, Shijie Huang, Fangtai Wu, Yubo Huang, Zhenxiong Tan, Songhua Liu, Jiaming Liu, Ruihua Huang
cs.AI

초록

고해상도 이미지 편집은 전문 작업 환경에서 점점 더 요구되고 있지만, 기존 확산 기반 모델은 이차 주의집중 복잡도와 과도한 메모리 요구량으로 인해 1K 미만의 해상도로 제한되어 있다. 널리 사용되는 해결책은 저해상도에서 편집한 후 독립적인 초해상도를 적용하는 2단계 파이프라인을 활용하는 것이다. 그러나 이 접근 방식은 두 가지 심각한 문제를 겪는다: 원본 고해상도(HR) 소스와 모순되는 환각 세부사항이 생성되는 정보 발산과, 과도하게 평활화되거나 과도하게 선명해진 아티팩트로 나타나는 질감 저하이다. 우리는 효율적인 초고해상도 편집을 위한 확산 브리지 프레임워크인 EditBridge를 제안한다. 노이즈로부터 재생성하는 기존 확산 방식과 달리, 우리는 정제 과정을 저해상도(LR) 편집 결과에서 그 고해상도 대응물로의 구조화된 데이터-데이터 변환으로 공식화하며, 원본의 세부사항을 보존하기 위해 원본 HR 소스에 명시적으로 조건화한다. HR 소스 안내를 효율적으로 통합하기 위해, 우리는 1단계 편집에서 얻은 의미적 대응 관계를 활용하여 교차 이미지 상호작용을 공간적으로 정렬된 영역으로 제한함으로써 계산 오버헤드를 크게 줄이는 사전 안내 블록 단위 희소 주의집중 메커니즘을 도입한다. 광범위한 실험을 통해 EditBridge가 최대 4K 해상도에서 우수한 지각 품질로 높은 충실도의 편집을 달성하고, 2K에서 3.6~8.4배의 속도 향상을 제공하며, 61초 만에 실용적인 4K 편집을 가능하게 함을 입증한다.
English
High-resolution image editing is increasingly demanded in professional workflows, yet existing diffusion-based models remain constrained to resolutions below 1K due to quadratic attention complexity and prohibitive memory requirements. A prevalent workaround employs a two-stage pipeline: editing at low resolution followed by independent super-resolution. However, this approach suffers from two critical issues: information divergence, where hallucinated details contradict the original high-resolution (HR) source, and texture degradation, manifesting as over-smoothed or over-sharpened artifacts. We propose EditBridge, a diffusion bridge framework for efficient ultra high-resolution editing. Unlike conventional diffusion that regenerates from noise, we formulate refinement as structured data-to-data translation from the low-resolution (LR) edited result to its HR counterpart, explicitly conditioned on the original HR source to preserve authentic details. To efficiently incorporate HR source guidance, we introduce a prior-guided block-wise sparse attention mechanism that exploits semantic correspondence from first-stage editing to constrain cross-image interactions to spatially aligned regions, significantly reducing computational overhead. Extensive experiments demonstrate that EditBridge achieves high-fidelity editing with superior perceptual quality at resolutions up to 4K, delivering 3.6--8.4times speedup at 2K and enabling practical 4K editing in 61 seconds.