ChatPaper.aiChatPaper

Poplar:一种可扩展的以人为中心的图像数据集合成流水线

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis

August 1, 2026
作者: Zhishan Zou
cs.AI

摘要

近期的图像生成器能够合成具有说服力的以人为中心的图像,然而生成一个有用的集合与生成单张成功图像是两回事。以人为中心的数据集必须覆盖多样化的人物与情境,避免不合理的属性组合,保留日常摄影的视觉特征,并在大规模尺度上公开质量控制决策。我们提出Poplar——一种可复现的“指定—渲染—审查”(Specify–Render–Inspect)流水线,用于以人为中心的图像数据集合成。“指定”阶段在常识约束下对结构化属性进行采样,并将其转化为面向摄影的提示词。“渲染”阶段采用经过真实感适配的图像生成器,结合构图感知的宽高比设置,并对明显的技术失败进行重试。“审查”阶段对每个候选图像执行单一的结构化视觉—语言评估,在保留原始提示词的同时,剔除内在图像缺陷或明显的提示词不匹配。利用Poplar,我们构建了Poplar-9K数据集:从11,765个经审查的候选中保留9,401对精选的以人为中心的图像—文本对(接受率79.9\%)。我们随数据集一并发布流水线、配置、不可更改的生成提示词以及可审计的审查记录,作为构建可定制以人为中心集合的紧凑资源。
English
Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains different from producing a single successful image. A human-centric dataset must cover varied people and contexts, avoid implausible attribute combinations, preserve an everyday photographic character, and expose quality-control decisions at scale. We present Poplar, a reproducible Specify--Render--Inspect pipeline for human-centric image dataset synthesis. Specify samples structured attributes under commonsense constraints and verbalizes them as photography-oriented prompts. Render uses a realism-adapted image generator across composition-aware aspect ratios and retries obvious technical failures. Inspect applies a single structured vision--language review to each candidate, preserving the original prompt while rejecting intrinsic image defects or material prompt mismatches. Using Poplar, we construct Poplar-9K: 9,401 curated human-centric image--text pairs retained from 11,765 reviewed candidates (79.9\% acceptance). We release the dataset together with the pipeline, configurations, immutable generation prompts, and auditable inspection records as a compact resource for building customizable human-centric collections.