Oxygen-TryOn: 모든 아이템 가상 피팅을 위한 패션 네이티브 파운데이션 모델
Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On
July 23, 2026
저자: Yong Liu, Xiaolong Fu, Zihang Xu, Wen Xue, Xueheng Li, Lin Song, Yuan Zhang, Chuyang Zhao, Haoyang Huang, Nan Duan, Yipeng Sun, Yan Li, Simiu Gu
cs.AI
초록
우리는 Oxygen-TryOn을 소개한다. 이는 모든 아이템을 대상으로 하는 가상 피팅을 위한 통합 기초 모델이다. 범용 이미지 편집기를 재활용하는 대신, Oxygen-TryOn은 패션 네이티브(fashion-native) 방식으로 설계되어 전용 데이터 엔진과 피팅 특화 학습을 통해 구축되었다. 하나 이상의 참조 아이템(깨끗한 제품 사진 또는 현장에서 촬영된 착용 사진)과 단일 대상 인물 이미지가 주어지면, 거의 모든 패션 카테고리에 걸쳐 대상이 해당 아이템을 착용한 사실적인 이미지를 합성한다. 기존 시스템은 스튜디오 환경에서 단일 의류 카테고리를 처리하며, 최근의 다중 참조 방법도 여전히 의류 중심에 머물러 있다. 반면, Oxygen-TryOn은 전신 및 반신 뷰, 가변적인 참조 수, 자유로운 다중 아이템 구성을 포함한 다양한 아이템과 시나리오를 지원하면서, 대상 인물의 정체성과 아이템 외관을 모두 충실히 보존한다. 마스크 기반 인페인팅 대신, 우리는 피팅을 다중 참조, 이해 기반 생성 작업으로 재정의한다. 대규모로 고품질 피팅 데이터를 수집, 제조, 주석 및 필터링하는 데이터 엔진을 구축하고, 지속적 사전 학습(CPT), 지도 미세 조정(SFT), 강화 학습(RL)의 3단계 레시피를 설계한다. RL 단계는 자체 피팅 보상 모델과 독점적인 루브릭 기반 범용 모델을 결합한 하이브리드 보상을 사용하여, 세부 일관성과 명령 수준 품질을 공동으로 감독한다. 또한 동일한 과정에서 일반 편집 명령(예: 포즈 변경)도 따르도록 한다. 공개 벤치마크와 자체 Oxygen-TryOn Bench에서 단일 아이템 피팅에 대해 최첨단 일관성과 사실성을 달성하고, 다중 아이템 피팅에서 선두를 차지하여 주요 독점 시스템(Nano Banana Pro, GPT-Image-2, Seedream5 Lite)과 오픈소스 모델(FLUX.2) 모두에 필적하거나 능가한다.
English
We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).