--- library_name: tpips base_model: Qwen/Qwen3-VL-Embedding-8B pipeline_tag: image-feature-extraction tags: - text-conditioned - perceptual-similarity - image-similarity - vision-language license: other --- # TPIPS — Embedding (late fusion) Text-conditioned perceptual image similarity, built on [Qwen/Qwen3-VL-Embedding-8B](https://huggingface.co/Qwen/Qwen3-VL-Embedding-8B). This repo holds the **`embedding`** checkpoint (one of three TPIPS models, each in its own repo — see the table at the bottom). Code and full docs: [https://github.com/adobe-research/TPIPS](https://github.com/adobe-research/TPIPS). **Late fusion.** Each `(text, image)` is encoded independently into an L2-normalised embedding. The pairwise score is `cos(e_a, e_b)` (higher = more similar). Odd-one-out probabilities are a `softmax` over the three "other-pair" scores divided by the temperature; 2AFC compares the two reference-candidate scores. The pairwise score is the model's raw output (temperature is applied at the probability step). | Property | Value | |----------|-------| | Base model | Qwen/Qwen3-VL-Embedding-8B | | Pairwise score | `cos(e_a, e_b)` | | Fine-tuning | LoRA (r=16, α=32) on the LLM layers | | Pooling | last-token | | Temperature | 0.05 (applied at the probability step) | | Prompt `X` | `Represent the similarity of the image based on X.` | ## Usage TPIPS supports Python 3.10 and later. Install matching PyTorch and torchvision builds from the [official PyTorch installer](https://pytorch.org/get-started/locally/), then install TPIPS: ```bash pip install tpips ``` Start with the recommended `embedding` model: ```python import tpips from PIL import Image model = tpips.load_model("embedding", device="cuda") a = Image.open("a.jpg").convert("RGB") b = Image.open("b.jpg").convert("RGB") similarity = model.similarity(a, b, factor="lighting") # higher is more similar distance = model.distance(a, b, factor="lighting") # lower is more similar ``` The first call downloads the selected TPIPS checkpoint and its Qwen backbone. A CUDA GPU is recommended; FlashAttention is optional. ## The TPIPS models | Model | Repo | |-------|------| | Embedding (late fusion) | [`sywang/TPIPS-Embed-Qwen3VL-8B`](https://huggingface.co/sywang/TPIPS-Embed-Qwen3VL-8B) | | Early Fusion | [`sywang/TPIPS-EarlyFusion-Qwen3VL-8B`](https://huggingface.co/sywang/TPIPS-EarlyFusion-Qwen3VL-8B) | | Activation Distance | [`sywang/TPIPS-ActDiff-Qwen3VL-8B`](https://huggingface.co/sywang/TPIPS-ActDiff-Qwen3VL-8B) | ## License TPIPS is provided under the [Adobe Research License](https://github.com/adobe-research/TPIPS/blob/main/LICENSE) for noncommercial research use. See the license for the complete terms.