Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

Yong Liu, Xiaolong Fu, Zihang Xu, Wen Xue, Xueheng Li, Lin Song, Yuan Zhang, Chuyang Zhao, Haoyang Huang, Nan Duan, Yipeng Sun, Yan Li, Simiu Gu

arXiv:2607.21694 · 2026-07-28 공개 · arXiv · PDF

reinforcement-learning foundation-model supervised-fine-tuning image-synthesis multi-reference data-engine virtual-try-on fashion-ai

Abstract

We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).

한국어 요약

한 줄 요약

Oxygen-TryOn은 다양한 의류와 액세서리를 포함한 항목을 정밀하게 시뮬레이션하는 통합 가상 시착 모델로, 기존 시스템보다 뛰어난 일관성과 사실감을 보인다.

핵심 기여도

핵심 아이디어

기존 시스템은 주로 단일 의류 카테고리에 국한되어 있었으며, 마스크 기반 인페인팅 방식으로 인해 세부 텍스처나 구조를 정확히 재현하지 못했다. Oxygen-TryOn은 이러한 문제를 해결하기 위해 **multi-reference, understanding-driven generation task**로 시착 문제를 재정의했다. 모델은 주어진 참조 항목(제품 이미지 또는 실제 착용 사진)을 이해하고, 항목 간의 레이어링, 가림, 변형을 고려하여 시착 이미지를 생성한다. 이는 단순히 정의된 영역을 채우는 것이 아니라, 항목의 구조와 착용 방식을 이해하는 과정을 포함한다. 이 접근법은 **Oxygen-TryOn**이 **unseen items**와 **complex compositions**에도 강한 일반화 능력을 갖도록 한다.

기술적 접근법

주요 결과

의의 및 한계

Oxygen-TryOn은 **any-item**, **multi-reference** 시착을 **high-fidelity**로 처리하는 최초의 시스템으로, 기존 모델들이 단일 카테고리나 제한된 환경에 의존했던 문제를 해결한다. 또한, **CPT–SFT–RL** 훈련 레시피와 **JoyAI-Image-Edit** 기반 모델을 통해 전문적인 시착 모델을 구축하는 투명하고 재현 가능한 방법을 제시한다. 그러나 **in-house try-on reward model**과 **rubric-guided general-purpose model**은 외부에서 접근하기 어려운 **proprietary systems**에 의존하므로, 완전한 오픈소스화는 아직 이루어지지 않았다. 또한, **Cloth-to-Model**과 **Model-to-Model** 시나리오 모두에서 **Usability Rate**는 80% 이하로, 일부 결과는 수작업 수정이 필요하다.

실용적 활용

Oxygen-TryOn은 온라인 쇼핑 플랫폼에서 사용자에게 의류, 신발, 액세서리 등을 시뮬레이션하여 실제 착용 효과를 제공할 수 있으며, 특히 **JINGDONG TryOn**과 같은 실제 산업 적용 사례가 존재한다. 또한, **multi-item composition**과 **complex background** 처리 능력을 갖추고 있어, 다양한 상황에서 활용 가능한 **virtual try-on** 솔루션으로 기능한다.