CtrlVTON: Experimentação Virtual Controlável via Segmentação Visual-Instância-Prompt
CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation
July 10, 2026
Autores: Seungyong Lee, Hyun Jun Jang, Sangoh Kim, Sungjoon Park
cs.AI
Resumo
A prova virtual (VTO) obteve progressos significativos na transferência realista de peças de vestuário para uma pessoa-alvo. No entanto, a maioria dos sistemas oferece ao utilizador pouco controlo sobre a forma como uma peça deve ser vestida – o seu tamanho (folgado ou justo), estilo (por exemplo, enfiado ou solto, aberto ou fechado) e posicionamento espacial no corpo. Abordamos esta lacuna com duas contribuições complementares. Primeiro, definimos e resolvemos a Segmentação por Prompt de Instância Visual via VIP-SAM: dada uma imagem flatlay de uma peça de vestuário, segmentar essa instância específica numa fotografia de uma pessoa a vesti-la. Esta é uma tarefa ao nível da instância, distinta da segmentação ao nível da categoria habitualmente estudada. Segundo, introduzimos o CtrlVTON, uma estrutura VTO controlável que reformula a prova virtual como um problema de edição de imagem e adiciona máscaras de segmentação como controlo ao nível do pixel sobre a disposição da peça, incluindo estilo, tamanho e posicionamento espacial no corpo. O VIP-SAM e o CtrlVTON alcançam, cada um, resultados de estado da arte nas respetivas tarefas. Em particular, o CtrlVTON gera imagens que seguem as disposições fornecidas pelo utilizador de forma muito mais fiel do que os sistemas de edição proprietários mais robustos, igualando-os em termos de fidelidade da peça.
English
Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the user little control over how a garment should be worn -- its size (loose or fitted), style (e.g., tucked in or untucked, open or closed), and spatial placement on the body. We address this gap with two complementary contributions. First, we define and solve Visual-Instance-Prompt Segmentation via VIP-SAM: given a flatlay image of a garment, segment that specific instance in a photograph of a person wearing it. This is an instance-level task, distinct from the typically studied category-level segmentation. Second, we introduce CtrlVTON, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body. VIP-SAM and CtrlVTON each achieve state-of-the-art results on their respective tasks. In particular, CtrlVTON generates images that follow user-provided layouts far more faithfully than the strongest proprietary editing systems while matching them on garment fidelity.