Instructions to use ArtmeScienceLab/Garments2Look-LoRA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use ArtmeScienceLab/Garments2Look-LoRA with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Qwen/Qwen-Image-Edit-2509", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("ArtmeScienceLab/Garments2Look-LoRA") prompt = "Turn this cat into a dog" input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png") image = pipe(image=input_image, prompt=prompt).images[0] - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("Qwen/Qwen-Image-Edit-2509", dtype=torch.bfloat16, device_map="cuda")
pipe.load_lora_weights("ArtmeScienceLab/Garments2Look-LoRA")
prompt = "Turn this cat into a dog"
input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png")
image = pipe(image=input_image, prompt=prompt).images[0]Garments2Look-LoRA
Task-specific LoRA adapters for Qwen-Image-Edit-2509, trained for outfit-level virtual try-on with multiple garments and accessories.
- Model repository: ArtmeScienceLab/Garments2Look-LoRA
- Code and training instructions: Garments2Look
- Dataset: ArtmeScienceLab/Garments2Look
- Paper: Garments2Look (CVPR 2026)
Release status: both task-specific LoRA checkpoints, training/inference code, and updated dataset inputs are publicly available. Use the code repository for installation and inference, and the dataset card for download and preparation.
The released adapters were trained further than the CVPR rebuttal-stage models. With additional training and more suitable inpainting masks, we expect improved results compared with the earlier models.
The updated dataset provides 98,012 outfit records with v1.1 annotations, OOTD collages, editing source images, and five annotation types: ATR, DensePose, DWPose, LIP, and refined v3 masks with dilation. For inpainting, prepare Figure 1 using the v3 mask; for editing, use the provided edited/banana/ source image. Dataset preparation instructions and split counts are maintained in the dataset card.
Checkpoints
| Task | Checkpoint | Training |
|---|---|---|
| Inpainting | Qwen-Image-Edit-2509-LoRA-2-refer-20k-inpainting-epoch-1.safetensors |
20K samples, 2 completed epochs, LoRA rank 32 |
| Editing | Qwen-Image-Edit-2509-LoRA-2-refer-20k-editing-epoch-1.safetensors |
20K samples, 2 completed epochs, LoRA rank 32 |
These files are LoRA adapters, not complete base models. epoch-1 is zero-indexed and denotes the checkpoint saved after the second epoch. Use the checkpoint matching your task.
Garments2Look-LoRA/
βββ README.md
βββ examples/comparison.jpg
βββ manifest.json
βββ Qwen-Image-Edit-2509-LoRA-2-refer-20k-inpainting-epoch-1.safetensors
βββ Qwen-Image-Edit-2509-LoRA-2-refer-20k-editing-epoch-1.safetensors
Inputs
Both tasks take two images and a text prompt:
| Input | Inpainting | Editing |
|---|---|---|
| Figure 1 | Person image with clothing regions masked in gray (128) | Source person image wearing an existing outfit |
| Figure 2 | OOTD collage of target reference items | OOTD collage of target reference items |
| Prompt | Numbered items, styling instructions, and layering order | Numbered items, styling instructions, and layering order |
Keep the collage item order consistent with the numbered prompt. The output is a person image wearing the target outfit. No separate mask argument is passed to inference: inpainting uses the already masked person image as Figure 1.
Installation and download
Run from the code repository root. The tested environment uses Python 3.10, PyTorch 2.7.1 with CUDA 12.8, and an NVIDIA H200.
git clone https://github.com/ArtmeScienceLab/Garments2Look.git
cd Garments2Look
conda create -n g2l-lora python=3.10 -y
conda activate g2l-lora
python -m pip install torch==2.7.1 torchvision==0.22.1 --index-url https://download.pytorch.org/whl/cu128
python -m pip install -r requirements.txt
python -m pip install -e . --no-deps
hf download Qwen/Qwen-Image-Edit-2509 --local-dir models/Qwen-Image-Edit-2509
# Download the released task-specific adapters.
hf download ArtmeScienceLab/Garments2Look-LoRA --local-dir models/Garments2Look-LoRA
Inpainting inference
python scripts/inference/inference.py --task inpainting \
--model-dir models/Qwen-Image-Edit-2509 \
--lora models/Garments2Look-LoRA/Qwen-Image-Edit-2509-LoRA-2-refer-20k-inpainting-epoch-1.safetensors \
--origin examples/input-inpainting.png --ootd examples/ootd.png \
--prompt-file examples/prompt-inpainting.txt \
--output output/inpainting.png --seed 123 --steps 40 --cfg-scale 4.0
Editing inference
python scripts/inference/inference.py --task editing \
--model-dir models/Qwen-Image-Edit-2509 \
--lora models/Garments2Look-LoRA/Qwen-Image-Edit-2509-LoRA-2-refer-20k-editing-epoch-1.safetensors \
--origin examples/input-editing.png --ootd examples/ootd.png \
--prompt-file examples/prompt-editing.txt \
--output output/editing.png --seed 123 --steps 40 --cfg-scale 4.0
The default inference seed is 123. Each run saves an output PNG and a JSON record containing its prompt and parameters. Images are aligned to multiples of 16 within a 1,048,576-pixel budget. Set CUDA_VISIBLE_DEVICES to select a GPU. For a custom outfit, replace the origin image, OOTD collage, and prompt file.
Use the existing server weights
The same commands work with absolute local paths; no adapter download is needed:
# Run from /home/hjy/repo/ours/Garments2Look/github-repo.
export BASE_MODEL=/mount/data/hjy/models/Qwen/Qwen-Image-Edit-2509
export LORA_DIR=/mount/data/hjy/models/Garments2Look-LoRA
python scripts/inference/inference.py --task inpainting \
--model-dir "$BASE_MODEL" --lora "$LORA_DIR/Qwen-Image-Edit-2509-LoRA-2-refer-20k-inpainting-epoch-1.safetensors" \
--origin examples/input-inpainting.png --ootd examples/ootd.png \
--prompt-file examples/prompt-inpainting.txt --output output/inpainting.png
python scripts/inference/inference.py --task editing \
--model-dir "$BASE_MODEL" --lora "$LORA_DIR/Qwen-Image-Edit-2509-LoRA-2-refer-20k-editing-epoch-1.safetensors" \
--origin examples/input-editing.png --ootd examples/ootd.png \
--prompt-file examples/prompt-editing.txt --output output/editing.png
Example
Columns: OOTD, inpainting input/output, editing input/output. The full prompts appear as two lines below the images. This six-item test example uses seed 123, 40 steps, and guidance 4.0 for both tasks.
Full prompt:
Keep the woman's identity, pose, background in Figure 1 unchanged, wearing the outfit in Figure 2, include (1) a top (partially unbuttoned, tucked-in), (2) a sweater (unbuttoned), (3) pants, (4) loafers, (5) a bag, (6) a belt (worn around waist). Layering Order: (1) -> (6) -> (2).
Training
The code repository includes data preparation and LoRA training for both tasks. Generate task-specific metadata from the training split, then run:
export MODEL_DIR="$PWD/models/Qwen-Image-Edit-2509"
export DATASET_ROOT=/path/to/Garments2Look-data
export TASK=inpainting
export METADATA="$PWD/data/metadata/train-inpainting.json"
export OUTPUT_DIR="$PWD/models/train/$TASK"
NPROC=1 SEED=123 bash scripts/train/train_lora.sh
# Editing: use editing-task metadata and a separate output directory.
export TASK=editing
export METADATA="$PWD/data/metadata/train-editing.json"
export OUTPUT_DIR="$PWD/models/train/$TASK"
NPROC=1 SEED=123 bash scripts/train/train_lora.sh
Defaults: rank 32, learning rate 1e-4, two epochs, gradient checkpointing, and a 1,048,576-pixel budget. For data preparation, multi-GPU usage, and metadata fields, see the code README. Editing-task source images are available in the updated dataset. The bundled examples are test samples and should not be used as a training benchmark.
Limitations and license
Fine accessory details, garment fidelity, styling, and pose preservation may vary. The example is illustrative, not an aggregate evaluation. The base model is required, and its license applies separately. An explicit adapter license has not yet been specified in this model repository.
Citation
If you use our dataset or models in your research, please consider citing our paper:
@inproceedings{cvpr2026garments2look,
title={Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories},
author={Hu, Junyao and Cheng, Zhongwei and Wong, Waikeung and Zou, Xingxing},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2026}
}
- Downloads last month
- 10
Model tree for ArtmeScienceLab/Garments2Look-LoRA
Base model
Qwen/Qwen-Image-Edit-2509
# Gated model: Login with a HF token with gated access permission hf auth login