CtrlVTON: 視覚インスタンスプロンプトセグメンテーションによる制御可能なバーチャル試着
CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation
July 10, 2026
著者: Seungyong Lee, Hyun Jun Jang, Sangoh Kim, Sungjoon Park
cs.AI
要旨
バーチャル試着(VTO)は、対象人物への衣類のリアルな転写において大きな進歩を遂げている。しかし、ほとんどのシステムでは、ユーザーが衣類の着用方法(サイズの緩さやフィット感、スタイル(例:インまたはアウト、開襟か閉襟)、身体上の空間的配置)をほとんど制御できない。本研究では、2つの相補的な貢献によりこのギャップに対処する。第一に、VIP-SAMを用いたVisual-Instance-Prompt Segmentationを定義し解決する。すなわち、衣類のフラットレイ画像を与えられた際に、それを着用した人物の写真から当該インスタンスをセグメント化する。これは、通常研究されるカテゴリレベルのセグメンテーションとは異なる、インスタンスレベルのタスクである。第二に、制御可能なVTOフレームワークであるCtrlVTONを導入する。これは試着を画像編集問題として捉え直し、セグメンテーションマスクを、スタイル、サイズ、身体上の空間的配置を含む衣類のレイアウトに対するピクセルレベルの制御として追加する。VIP-SAMとCtrlVTONは、それぞれのタスクで最先端の成果を達成している。特にCtrlVTONは、最も強力なプロプライエタリ編集システムと同等の衣類忠実度を維持しつつ、ユーザーが提供するレイアウトをはるかに忠実に再現した画像を生成する。
English
Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the user little control over how a garment should be worn -- its size (loose or fitted), style (e.g., tucked in or untucked, open or closed), and spatial placement on the body. We address this gap with two complementary contributions. First, we define and solve Visual-Instance-Prompt Segmentation via VIP-SAM: given a flatlay image of a garment, segment that specific instance in a photograph of a person wearing it. This is an instance-level task, distinct from the typically studied category-level segmentation. Second, we introduce CtrlVTON, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body. VIP-SAM and CtrlVTON each achieve state-of-the-art results on their respective tasks. In particular, CtrlVTON generates images that follow user-provided layouts far more faithfully than the strongest proprietary editing systems while matching them on garment fidelity.