Oxygen-TryOn: ファッション特化型基盤モデルによる任意アイテムのバーチャル試着
Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On
July 23, 2026
著者: Yong Liu, Xiaolong Fu, Zihang Xu, Wen Xue, Xueheng Li, Lin Song, Yuan Zhang, Chuyang Zhao, Haoyang Huang, Nan Duan, Yipeng Sun, Yan Li, Simiu Gu
cs.AI
要旨
我々は、任意アイテムのバーチャル試着のための統合基盤モデル「Oxygen-TryOn」を提案する。汎用画像編集ツールを流用するのではなく、Oxygen-TryOnはファッションに特化したネイティブなモデルであり、専用のデータエンジンと試着特化の学習により構築されている。一枚または複数の参照アイテム(清潔な商品画像または実環境での着用写真)と、単一の対象人物画像を入力として、ほぼあらゆるファッションカテゴリにおいて、対象者がそのアイテムを着用したフォトリアリスティックな画像を合成する。従来のシステムは、スタジオ環境で単一の衣料カテゴリを扱うものであり、最近のマルチ参照手法も依然として衣料中心である。対照的にOxygen-TryOnは、全身・半身ビュー、可変数の参照、自由なマルチアイテム合成を含む多様なアイテムとシナリオをサポートし、対象者のアイデンティティとアイテムの外観の両方を忠実に保持する。マスクベースのインペインティングではなく、試着をマルチ参照・理解駆動型の生成タスクとして再定義する。我々は、高品質な試着データを大規模に収集・生成・アノテーション・フィルタリングするデータエンジンを構築し、継続事前学習(CPT)、教師ありファインチューニング(SFT)、強化学習(RL)の3段階のレシピを設計する。RL段階では、社内の試着リワードモデルと、独自のルーブリックに基づく汎用モデルを組み合わせたハイブリッドリワードを用い、細粒度の一貫性と指示レベルの品質を共同で監視する。また、同じパスで一般的な編集指示(ポーズ変更など)にも従う。公開ベンチマークおよび社内のOxygen-TryOn Benchにおいて、単一アイテム試着で最先端の一貫性と写実性を達成し、マルチアイテム試着でもリードしており、主要なプロプライエタリシステム(Nano Banana Pro、GPT-Image-2、Seedream5 Lite)およびオープンソースモデル(FLUX.2)に匹敵または凌駕する。
English
We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).