ChatPaper.aiChatPaper

次に何を編集するか:対話システムにおける視覚的に整合された画像編集フォローアップ提案

What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems

August 3, 2026
著者: Zhijing Zhang, Jinpeng Yu, Xin Song, Bingnan Li, Chuyue Li, Changhui Du, Xiaolin Fang, Jiaming Liu, Ruihua Huang
cs.AI

要旨

対話型アシスタントは、ユーザーがタスクを継続できるようフォローアップ編集を推奨することが増えている。既存のシステムは主にテキストのみの対話を対象としており、画像生成対話は十分に検討されていない。画像生成タスクにおいて、有用なフォローアップ編集提案は、ユーザーの好みを反映し、多様な方向性を提供し、現在の画像に対して実行可能でなければならない。我々はQwenアプリから10万件の実マルチターン画像生成対話サンプルを収集し、その80.1%が画像に依存していることを発見し、マルチモーダルレコメンデーションの必要性を裏付けた。本研究では、この設定に3段階のフレームワークで取り組む。段階1では、実オンラインデータを用いて、人間がレビューした適切なフォローアップ編集意図のテーブルを構築し、SFTターゲットを作成してマルチモーダルポリシーをファインチューニングする。段階2では、ルール誘導型SFT提案を実際のユーザー選択と整合させるため、ユーザーのクリックフィードバックを用いて多目的強化学習によりポリシーを最適化する。段階3では、提案された編集と現在の画像との間の視覚的不整合を低減するため、追加のトレーニング監視として視覚検証器を導入する。広範な実験により、本フレームワークが自動評価と人間評価の両方においてベースラインを大幅に上回ることが実証された。数百万ユーザーを対象とした実環境でのユーザー無作為化A/Bテストでは、最終フレームワークにより視覚的不整合が3.7%から0.9%に低減された。さらに、レコメンデーションCTRが32.70%、画像持ち帰り率が16.32%、ユーザー1人当たりの平均対話ターン数が39.90%大幅に向上した(すべてp<0.05)。
English
Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only interactions, leaving image-creation conversations underexplored. In image-creation tasks, useful follow-up edit suggestions must reflect user preferences, offer diverse directions, and remain executable on the current image. We collected 100,000 real multi-turn image-creation conversation samples from Qwen App and found that 80.1% are image-dependent, underscoring the need for multimodal recommendation. We address this setting with a three-stage framework. In Stage 1, we use real online data to build a human-reviewed table of appropriate follow-up editing intents, then create SFT targets and fine-tune a multimodal policy. In Stage 2, to align rule-guided SFT suggestions with actual user choices, we use user click feedback to optimize the policy through multi-objective reinforcement learning. In Stage 3, to reduce visual inconsistencies between suggested edits and the current image, we introduce a visual verifier as additional training supervision. Extensive experiments demonstrate that our framework significantly outperforms baselines on both automatic and human evaluations. In a live user-randomized A/B test with millions of users, our final framework reduces visual inconsistency from 3.7% to 0.9%. Furthermore, it significantly improves recommendation CTR by 32.70%, image take-away rate by 16.32%, and average conversation turns per user by 39.90% (all p<0.05).