ChatPaper.aiChatPaper

Oxygen-TryOn:面向任意物品虚拟试穿的时尚原生基础模型

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

July 23, 2026
作者: Yong Liu, Xiaolong Fu, Zihang Xu, Wen Xue, Xueheng Li, Lin Song, Yuan Zhang, Chuyang Zhao, Haoyang Huang, Nan Duan, Yipeng Sun, Yan Li, Simiu Gu
cs.AI

摘要

我们提出Oxygen-TryOn,一个面向任意商品虚拟试穿的统一基础模型。与改造通用图像编辑器不同,Oxygen-TryOn是专为时尚原生设计的模型,通过专用数据引擎和针对试穿任务的训练构建。给定一个或多个参考商品(清晰的实物图或真实场景中的穿着照片)以及单张目标人物图像,它能合成人物穿着这些商品的照片级真实图像,覆盖几乎任何时尚品类。现有系统在影棚环境下处理单一服装品类,而近期的多参考方法仍以服装为中心;相比之下,Oxygen-TryOn支持多样化的商品和场景,包括全身和半身视图、可变数量的参考图以及自由的多商品组合,同时忠实保留人物身份和商品外观。我们不再采用基于掩码的修复方法,而是将试穿重新定义为多参考、理解驱动的生成任务。我们构建了一个数据引擎,能够大规模收集、制造、标注和筛选高质量试穿数据,并设计了持续预训练(CPT)、监督微调(SFT)和强化学习(RL)三阶段方案。RL阶段采用混合奖励,结合内部试穿奖励模型和专有的、基于评分指南的通用模型,共同监督细粒度一致性和指令级质量。它还能在同一流程中遵循通用编辑指令(如姿势变化)。在公开基准测试和我们内部的Oxygen-TryOn Bench上,它在单商品试穿上实现了最先进的一致性和真实感,在多商品试穿上领先,匹配或超越了领先的专有系统(Nano Banana Pro、GPT-Image-2、Seedream5 Lite)和开源模型(FLUX.2)。
English
We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).