다음에 무엇을 편집할까: 대화형 시스템에서의 시각적으로 정렬된 이미지 편집 후속 제안
What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
August 3, 2026
저자: Zhijing Zhang, Jinpeng Yu, Xin Song, Bingnan Li, Chuyue Li, Changhui Du, Xiaolin Fang, Jiaming Liu, Ruihua Huang
cs.AI
초록
대화형 어시스턴트는 사용자가 작업을 계속 이어갈 수 있도록 후속 편집안을 점점 더 많이 추천하고 있다. 기존 시스템은 주로 텍스트 전용 상호작용을 대상으로 하여 이미지 생성 대화는 충분히 탐구되지 않았다. 이미지 생성 작업에서 유용한 후속 편집 제안은 사용자 선호도를 반영하고, 다양한 방향을 제시하며, 현재 이미지에 대해 실행 가능해야 한다. 우리는 Qwen App에서 10만 개의 실제 다중 턴 이미지 생성 대화 샘플을 수집했으며, 이 중 80.1%가 이미지 의존적임을 발견하여 멀티모달 추천의 필요성을 확인했다. 우리는 이 문제를 3단계 프레임워크로 해결한다. 1단계에서는 실제 온라인 데이터를 사용하여 적절한 후속 편집 의도에 대한 사람이 검토한 테이블을 구축하고, SFT 타깃을 만들어 멀티모달 정책을 미세 조정한다. 2단계에서는 규칙 기반 SFT 제안을 실제 사용자 선택과 정렬하기 위해 사용자 클릭 피드백을 사용하여 다중 목표 강화 학습으로 정책을 최적화한다. 3단계에서는 제안된 편집과 현재 이미지 간의 시각적 불일치를 줄이기 위해 시각 검증기를 추가 훈련 감독으로 도입한다. 광범위한 실험을 통해 우리 프레임워크가 자동 평가와 인간 평가 모두에서 베이스라인을 유의미하게 능가함을 입증한다. 수백만 명의 사용자를 대상으로 한 실시간 사용자 무작위 배정 A/B 테스트에서 최종 프레임워크는 시각적 불일치를 3.7%에서 0.9%로 줄였다. 또한 추천 CTR(클릭률)을 32.70%, 이미지 가져가기 비율을 16.32%, 사용자당 평균 대화 턴 수를 39.90% 유의미하게 개선했다(모두 p<0.05).
English
Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only interactions, leaving image-creation conversations underexplored. In image-creation tasks, useful follow-up edit suggestions must reflect user preferences, offer diverse directions, and remain executable on the current image. We collected 100,000 real multi-turn image-creation conversation samples from Qwen App and found that 80.1% are image-dependent, underscoring the need for multimodal recommendation. We address this setting with a three-stage framework. In Stage 1, we use real online data to build a human-reviewed table of appropriate follow-up editing intents, then create SFT targets and fine-tune a multimodal policy. In Stage 2, to align rule-guided SFT suggestions with actual user choices, we use user click feedback to optimize the policy through multi-objective reinforcement learning. In Stage 3, to reduce visual inconsistencies between suggested edits and the current image, we introduce a visual verifier as additional training supervision. Extensive experiments demonstrate that our framework significantly outperforms baselines on both automatic and human evaluations. In a live user-randomized A/B test with millions of users, our final framework reduces visual inconsistency from 3.7% to 0.9%. Furthermore, it significantly improves recommendation CTR by 32.70%, image take-away rate by 16.32%, and average conversation turns per user by 39.90% (all p<0.05).