ChatPaper.aiChatPaper

Oxygen-TryOn:適用於任意商品虛擬試穿的時尚原生基礎模型

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

July 23, 2026
作者: Yong Liu, Xiaolong Fu, Zihang Xu, Wen Xue, Xueheng Li, Lin Song, Yuan Zhang, Chuyang Zhao, Haoyang Huang, Nan Duan, Yipeng Sun, Yan Li, Simiu Gu
cs.AI

摘要

我們提出 Oxygen-TryOn,一個針對任意品項的虛擬試穿統一基礎模型。不同於將通用型影像編輯器改作他用,Oxygen-TryOn 是專為時尚而生,透過專用資料引擎及試穿專屬訓練所打造。給定一個或多個參考品項(乾淨的商品照或非棚拍的穿著照)以及一張目標人物影像,它能合成出該人物穿著這些品項、且橫跨幾乎所有時尚類別的逼真影像。先前的系統僅能在棚拍環境中處理單一衣物類別,而近期的多參考方法仍以衣物為中心;相較之下,Oxygen-TryOn 支援多樣品項與場景,包括全身與半身視角、可變數量的參考,以及自由的多品項組合,同時忠實保留人物身分與品項外觀。我們不以遮罩式修補為基礎,而是將試穿重新定義為一項多參考、以理解為導向的生成任務。我們建立一個能大規模收集、製造、標註並篩選高品質試穿資料的資料引擎,並設計一套三階段的訓練配方:持續預訓練(CPT)、監督式微調(SFT)以及強化學習(RL)。RL 階段採用混合獎勵,結合內部試穿獎勵模型與專有、以評分指引為基礎的通用模型,共同監督細粒度一致性與指令層級品質。在同一流程中,它也能遵循一般的編輯指令(例如姿勢變化)。在公開基準測試與我們內部的 Oxygen-TryOn Bench 上,它在單品項試穿上達到最先進的一致度與真實感,並在多品項試穿上領先,能匹配或超越領先的專有系統(Nano Banana Pro、GPT-Image-2、Seedream5 Lite)與開源模型(FLUX.2)。
English
We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).