Poplar: 人間中心の画像データセット合成のためのスケーラブルなパイプライン
Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis
August 1, 2026
著者: Zhishan Zou
cs.AI
要旨
近年の画像生成器は、人間中心の説得力のある画像を合成できるが、有用なコレクションを生成することは、単一の成功した画像を生成することとは依然として異なる。人間中心のデータセットは、多様な人物と文脈をカバーし、非現実的な属性の組み合わせを避け、日常的な写真らしさを保ち、大規模な品質管理の判断を明示しなければならない。我々は、人間中心の画像データセット合成のための再現可能なSpecify--Render--Inspect(指定・レンダリング・検査)パイプラインであるPoplarを提案する。Specifyは、常識的な制約の下で構造化属性をサンプリングし、それらを写真指向のプロンプトとして言語化する。Renderは、構図を考慮したアスペクト比にわたって現実性に適応した画像生成器を使用し、明らかな技術的失敗を再試行する。Inspectは、各候補に単一の構造化された視覚・言語レビューを適用し、元のプロンプトを保持しながら、画像固有の欠陥または重大なプロンプト不一致を拒否する。Poplarを用いて、我々はPoplar-9Kを構築する:レビュー済み候補11,765件から保持された9,401件のキュレーション済み人間中心の画像・テキストペア(受理率79.9\%)。我々は、カスタマイズ可能な人間中心のコレクションを構築するためのコンパクトなリソースとして、パイプライン、構成、不変の生成プロンプト、および監査可能な検査記録とともにデータセットを公開する。
English
Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains different from producing a single successful image. A human-centric dataset must cover varied people and contexts, avoid implausible attribute combinations, preserve an everyday photographic character, and expose quality-control decisions at scale. We present Poplar, a reproducible Specify--Render--Inspect pipeline for human-centric image dataset synthesis. Specify samples structured attributes under commonsense constraints and verbalizes them as photography-oriented prompts. Render uses a realism-adapted image generator across composition-aware aspect ratios and retries obvious technical failures. Inspect applies a single structured vision--language review to each candidate, preserving the original prompt while rejecting intrinsic image defects or material prompt mismatches. Using Poplar, we construct Poplar-9K: 9,401 curated human-centric image--text pairs retained from 11,765 reviewed candidates (79.9\% acceptance). We release the dataset together with the pipeline, configurations, immutable generation prompts, and auditable inspection records as a compact resource for building customizable human-centric collections.