ChatPaper.aiChatPaper

超越《星夜》:捷徑感知控制狀態規劃於藝術家根基之文字轉圖像生成

Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation

August 7, 2026
作者: Kuan Xing, Ye Wang, Changyi Gan, Yuheng Li, Thao Nguyen, Yi Chang, Yilin Wang
cs.AI

摘要

藝術家導向的圖像生成,遠不止於在提示詞中附加藝術家姓名。圖像模型往往透過典型捷徑來回應藝術家姓名,例如重複出現的母題、通用色調,或過度表徵的時期特徵,而非保留使用者意圖中的場景。為此,我們提出 Atelier——一個用於藝術家導向圖像生成的捷徑感知控制狀態規劃框架。Atelier 將未充分明確的藝術意圖轉化為明確的控制狀態,其中區分場景錨點、保留/轉換決策、風格體系假設、角色綁定的藝術家證據,以及捷徑迴避約束。該框架以藝術家層級知識與局部圖塊參照為此狀態建立依據,編製後端感知的生成計畫,並透過全域與局部真實性回饋迭代地精煉候選結果。我們進一步提出 ArtIntentBench 基準,涵蓋梵谷與齊白石,包含藝術作品重繪、時期/風格控制生成、歷史未見主題、捷徑稽核,以及人類偏好評估。在開放權重與閉源生成器中,相較於提示工程、檢索增強與通用代理基線,Atelier 提升了藝術家層級的風格保真度,更忠實地保留來源結構,並大幅減少捷徑替換。這些結果表明,藝術家導向生成的瓶頸不僅在於圖像合成,更在於對明確且有證據支撐的藝術控制之上游推論。
English
Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical shortcuts, such as recurring motifs, generic palettes, or overrepresented period signatures, rather than preserving the user's intended scene. We introduce Atelier, a shortcut-aware control-state planning framework for artist-grounded image generation. Atelier translates underspecified artistic intent into an explicit control state that separates scene anchors, preserve/transform decisions, style-regime hypotheses, role-bound artist evidence, and shortcut-avoidance constraints. It grounds this state using artist-level knowledge and local patch references, compiles backend-aware generation plans, and iteratively refines candidates through global and local authenticity feedback. We further introduce ArtIntentBench, a benchmark covering Van Gogh and Qi Baishi across artwork re-rendering, period/style-controlled generation, historically unseen subjects, shortcut auditing, and human preference evaluation. Across open-weight and closed-source generators, Atelier improves artist-level style fidelity, preserves source structure more faithfully, and substantially reduces shortcut substitution compared with prompt-engineered, retrieval-augmented, and general-purpose agent baselines. These results suggest that artist-grounded generation is bottlenecked not only by image synthesis, but by the upstream inference of explicit, evidence-grounded artistic controls.