ChatPaper.aiChatPaper

下一步該編輯什麼:對話系統中視覺對齊的圖像編輯後續建議

What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems

August 3, 2026
作者: Zhijing Zhang, Jinpeng Yu, Xin Song, Bingnan Li, Chuyue Li, Changhui Du, Xiaolin Fang, Jiaming Liu, Ruihua Huang
cs.AI

摘要

對話式助手越來越常推薦後續編輯,以協助使用者延續任務。現有系統主要針對純文字互動,導致圖像生成對話領域仍未被充分探索。在圖像生成任務中,實用的後續編輯建議必須反映使用者偏好、提供多元方向,並且能在當前圖像上執行。我們從 Qwen App 收集了 100,000 個真實的多輪圖像生成對話樣本,發現其中 80.1% 具有圖像依賴性,凸顯了多模態推薦的必要性。我們以三階段框架來應對此情境。第一階段,我們利用真實的線上資料建立一份經人工審查的合適後續編輯意圖表,接著建構 SFT 目標並微調多模態策略。第二階段,為了使規則引導的 SFT 建議與實際使用者選擇一致,我們運用使用者點擊回饋,透過多目標強化學習最佳化策略。第三階段,為減少建議編輯與當前圖像之間的視覺不一致,我們引入視覺驗證器作為額外的訓練監督。大量實驗證明,我們提出的框架在自動評估與人工評估中均顯著優於基線方法。在涵蓋數百萬使用者的線上隨機 A/B 測試中,我們的最終框架將視覺不一致率從 3.7% 降至 0.9%。此外,推薦點擊率提升了 32.70%、圖片帶走率提升了 16.32%、每位使用者的平均對話輪數提升了 39.90%(全部 p<0.05)。
English
Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only interactions, leaving image-creation conversations underexplored. In image-creation tasks, useful follow-up edit suggestions must reflect user preferences, offer diverse directions, and remain executable on the current image. We collected 100,000 real multi-turn image-creation conversation samples from Qwen App and found that 80.1% are image-dependent, underscoring the need for multimodal recommendation. We address this setting with a three-stage framework. In Stage 1, we use real online data to build a human-reviewed table of appropriate follow-up editing intents, then create SFT targets and fine-tune a multimodal policy. In Stage 2, to align rule-guided SFT suggestions with actual user choices, we use user click feedback to optimize the policy through multi-objective reinforcement learning. In Stage 3, to reduce visual inconsistencies between suggested edits and the current image, we introduce a visual verifier as additional training supervision. Extensive experiments demonstrate that our framework significantly outperforms baselines on both automatic and human evaluations. In a live user-randomized A/B test with millions of users, our final framework reduces visual inconsistency from 3.7% to 0.9%. Furthermore, it significantly improves recommendation CTR by 32.70%, image take-away rate by 16.32%, and average conversation turns per user by 39.90% (all p<0.05).