CtrlVTON:基於視覺實例提示分割的可控虛擬試穿
CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation
July 10, 2026
作者: Seungyong Lee, Hyun Jun Jang, Sangoh Kim, Sungjoon Park
cs.AI
摘要
虚拟试穿(VTO)在将服装逼真地迁移至目标人物方面取得了显著进展。然而,大多数系统对用户如何穿着服装的控制十分有限——包括其尺寸(宽松或贴身)、款式(如扎入或外露、敞开或闭合)以及身体上的空间位置。我们通过两项互补性贡献填补了这一空白。首先,我们定义并解决了基于VIP-SAM的视觉实例提示分割问题:给定服装的平铺图像,从穿着该服装的人物照片中分割出该特定实例。这是一项实例级任务,与通常研究的类别级分割截然不同。其次,我们引入了CtrlVTON,一个可控的VTO框架,它将试穿重新定义为图像编辑问题,并添加分割掩码作为像素级控制,以调整服装布局,包括风格、尺寸以及身体上的空间位置。VIP-SAM和CtrlVTON各自在其任务上取得了最先进的结果。特别是,CtrlVTON生成的图像在遵循用户提供的布局方面远胜于最强大的专有编辑系统,同时与之在服装保真度上不相上下。
English
Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the user little control over how a garment should be worn -- its size (loose or fitted), style (e.g., tucked in or untucked, open or closed), and spatial placement on the body. We address this gap with two complementary contributions. First, we define and solve Visual-Instance-Prompt Segmentation via VIP-SAM: given a flatlay image of a garment, segment that specific instance in a photograph of a person wearing it. This is an instance-level task, distinct from the typically studied category-level segmentation. Second, we introduce CtrlVTON, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body. VIP-SAM and CtrlVTON each achieve state-of-the-art results on their respective tasks. In particular, CtrlVTON generates images that follow user-provided layouts far more faithfully than the strongest proprietary editing systems while matching them on garment fidelity.