利用属性引导的体裁扩展实现超越以故事为中心数据的创意写作规模化
Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion
August 14, 2026
作者: Hwan Chang, Yongil Kim, Heuiyeen Yeen, Yireun Kim, Jinsik Lee, Hwanhee Lee
cs.AI
摘要
面向大语言模型(LLM)的高质量创意写作数据仍主要以故事为中心,这限制了模型遵循多样化创意格式的结构与功能惯例的能力。我们提出了一种属性引导的体裁扩展框架,用于将创意写作数据扩展到故事生成之外。通过将主题广度与体裁形式控制相分离,该框架利用人工撰写的故事提示作为多样化的创意种子,同时借助人工整理的体裁属性来强制落实独特的结构、风格和格式规范。我们将这些要素结合起来,提示强大的大语言模型生成符合体裁的查询-响应对,并对其进行质量过滤。应用该框架,我们构建了多体裁集合(Multi-Genre Collection),这是一个包含5万个样本、涵盖13种创意体裁的语料库,包括故事、说唱、歌词、剧本、游戏设计、角色设计及其他创意格式。在分布外写作基准和留出体裁诊断上的实验表明,基于我们数据微调的模型不仅持续超越基础模型和写作专用基线,还超越了在现有写作语料库上训练的模型。体裁数量消融实验进一步表明,受控的体裁扩展——而非仅以故事为中心的数据扩展——是获得稳健创意写作能力的关键驱动因素。
English
High-quality creative writing data for large language models (LLMs) remains dominated by story-centric data, limiting models' ability to follow the structural and functional conventions of diverse creative formats. We propose an attribute-guided genre expansion framework for scaling creative writing data beyond story generation. By separating thematic breadth from genre-form control, our framework leverages human-authored story prompts as diverse creative seeds, while utilizing manually curated genre attributes to enforce distinct structural, stylistic, and formatting conventions. We combine these to prompt strong LLMs for genre-faithful query-response pairs, which are then quality-filtered. Applying this framework, we construct the Multi-Genre Collection, a 50K-example corpus spanning 13 creative genres, including story, rap, lyrics, scripts, game design, character design, and other creative formats. Experiments across out-of-distribution writing benchmarks and held-out genre diagnostics demonstrate that models fine-tuned on our data consistently surpass not only base models and writing-specialized baselines, but also models trained on existing writing corpora. Genre-count ablations further indicate that controlled genre expansion, rather than story-centric scaling alone, is a key driver of robust creative writing capability.