CtrlVTON:基于视觉实例提示分割的可控虚拟试穿
CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation
July 10, 2026
作者: Seungyong Lee, Hyun Jun Jang, Sangoh Kim, Sungjoon Park
cs.AI
摘要
虚拟试穿(VTO)在将服装真实地迁移到目标人物身上取得了显著进展。然而,大多数系统几乎没有让用户控制服装的穿着方式——其尺码(宽松或合身)、款式(例如,塞入或外露、敞开或闭合)以及身体上的空间位置。我们通过两项互补贡献来填补这一空白。首先,我们通过VIP-SAM定义并解决了视觉实例提示分割(Visual-Instance-Prompt Segmentation)任务:给定一件服装的平铺图,在穿着该服装的人物照片中分割出该特定实例。这是一项实例级任务,区别于通常研究的类别级分割。其次,我们提出了CtrlVTON,这是一个可控的VTO框架,它将试穿重新定义为图像编辑问题,并添加分割掩码作为服装布局的像素级控制,包括款式、尺码和身体上的空间位置。VIP-SAM和CtrlVTON分别在各自任务上达到了最先进的结果。特别地,CtrlVTON生成的图像在遵循用户提供的布局方面远强于最强专有编辑系统,同时在服装保真度上与它们相匹配。
English
Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the user little control over how a garment should be worn -- its size (loose or fitted), style (e.g., tucked in or untucked, open or closed), and spatial placement on the body. We address this gap with two complementary contributions. First, we define and solve Visual-Instance-Prompt Segmentation via VIP-SAM: given a flatlay image of a garment, segment that specific instance in a photograph of a person wearing it. This is an instance-level task, distinct from the typically studied category-level segmentation. Second, we introduce CtrlVTON, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body. VIP-SAM and CtrlVTON each achieve state-of-the-art results on their respective tasks. In particular, CtrlVTON generates images that follow user-provided layouts far more faithfully than the strongest proprietary editing systems while matching them on garment fidelity.