ChatPaper.aiChatPaper

Poplar:一種用於以人為中心的圖像數據集合成的可擴展管線

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis

August 1, 2026
作者: Zhishan Zou
cs.AI

摘要

近期的影像生成器能夠合成令人信服的以人為中心影像,但產出一個有用的集合不同於產出單張成功的影像。以人為中心的資料集必須涵蓋多元的人物與情境、避免不合理的屬性組合、保留日常攝影的特質,並在大規模下呈現品質控制的決策。我們提出 Poplar,一個可重現的「指定—渲染—檢查」流程,用於以人為中心的影像資料集合成。指定階段在常識約束下取樣結構化屬性,並將其表述為以攝影為導向的提示詞。渲染階段使用適應真實感的影像生成器,配合考量構圖的長寬比,並重試明顯的技術失敗。檢查階段對每個候選影像套用單一結構化視覺—語言審查,保留原始提示詞,同時拒絕內在影像缺陷或實質提示詞不符的情況。利用 Poplar,我們建構了 Poplar-9K:從 11,765 個經審查的候選影像中保留 9,401 個精選的以人為中心影像—文字對(接受率 79.9%)。我們同時釋出資料集與流程、設定、不可變的生成提示詞,以及可稽核的檢查記錄,作為建構可自訂以人為中心集合的精簡資源。
English
Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains different from producing a single successful image. A human-centric dataset must cover varied people and contexts, avoid implausible attribute combinations, preserve an everyday photographic character, and expose quality-control decisions at scale. We present Poplar, a reproducible Specify--Render--Inspect pipeline for human-centric image dataset synthesis. Specify samples structured attributes under commonsense constraints and verbalizes them as photography-oriented prompts. Render uses a realism-adapted image generator across composition-aware aspect ratios and retries obvious technical failures. Inspect applies a single structured vision--language review to each candidate, preserving the original prompt while rejecting intrinsic image defects or material prompt mismatches. Using Poplar, we construct Poplar-9K: 9,401 curated human-centric image--text pairs retained from 11,765 reviewed candidates (79.9\% acceptance). We release the dataset together with the pipeline, configurations, immutable generation prompts, and auditable inspection records as a compact resource for building customizable human-centric collections.