ChatPaper.aiChatPaper

평가-검증 보상 기반 일관된 다중 참조 이미지 편집

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

July 31, 2026
저자: Yingmao Miao, Pengfei Zhang, Xiaochen Lv, Meng Yu, Lei Sun, Xiangxiang Chu, Chao Shen, Chenhao Lin
cs.AI

초록

최근 이미지 편집 모델들이 빠른 발전을 이루었지만, 다중 참조 편집은 여전히 어려운 과제로 남아 있으며, 특히 참조 간 시각적 일관성 유지와 전반적인 시각적 조화 보장에서 그러하다. 강화 학습은 텍스트-이미지 생성 및 단일 이미지 편집에서 매우 효과적인 것으로 입증되었으나, 다중 참조 편집으로의 확장은 다중 이미지 관계 제약을 포착하는 적절한 보상 모델의 부재로 인해 제약을 받는다. 또한, 다중 모달 대규모 언어 모델(MLLM)을 제로샷 평가기로 단순 활용하는 방식은 환각에 취약한 장문 추론과 연역 능력이 제한적인 단문 판단 사이의 핵심적 긴장 관계에 직면한다. 본 연구에서는 이러한 문제를 다차원 평가-검증 보상(EVR)으로 해결한다. EVR은 평가를 개별적 시각 기준으로 분해하며, 각 기준에 대해 MLLM 평가기가 여러 후보 가설을 생성하고, 검증기가 각 주장을 구체적인 시각적 증거에 근거하여 수용 또는 기각함으로써 신뢰할 수 있고 세분화된 보상 신호를 생성한다. 확장 가능한 데이터 파이프라인과 함께, 본 방법은 구조 변경 없이 기성 편집 모델의 RL 파인튜닝을 가능하게 한다. 광범위한 실험을 통해 기본 Qwen-Image-Edit 대비 상당한 성능 향상을 보여주며, 일관성과 조화 측면에서 NanoBanana와 동등하거나 이를 능가하는 결과를 달성한다.
English
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has proven highly effective for text-to-image generation and single-image editing, but its extension to multi-reference editing is hindered by the absence of suitable reward models that capture multi-image relational constraints. Moreover, naively using multimodal large language models(MLLMs) as zero-shot evaluators faces a key tension between hallucination-prone long-form reasoning and the limited deductive power of short-form judgments. We address these issues with a Multi-dimensional Evaluation-Verification Reward(EVR). EVR decomposes evaluation into distinct visual criteria; for each criterion, an MLLM Evaluator generates multiple candidate hypotheses, and a Verifier grounds each claim in concrete visual evidence to accept or reject it, producing reliable and fine-grained reward signals. Together with a scalable data pipeline, our method enables RL fine-tuning of off-the-shelf editors without architectural changes. Extensive experiments show substantial gains over the base Qwen-Image-Edit, improving consistency and harmony to match or surpass NanoBanana.