| --- |
| library_name: tpips |
| base_model: Qwen/Qwen3-VL-Embedding-8B |
| pipeline_tag: image-feature-extraction |
| tags: |
| - text-conditioned |
| - perceptual-similarity |
| - image-similarity |
| - vision-language |
| license: other |
| --- |
| |
| # TPIPS — Embedding (late fusion) |
|
|
| Text-conditioned perceptual image similarity, built on |
| [Qwen/Qwen3-VL-Embedding-8B](https://huggingface.co/Qwen/Qwen3-VL-Embedding-8B). This repo holds the |
| **`embedding`** checkpoint (one of three TPIPS models, each in its own repo — see |
| the table at the bottom). Code and full docs: [https://github.com/adobe-research/TPIPS](https://github.com/adobe-research/TPIPS). |
|
|
| **Late fusion.** Each `(text, image)` is encoded independently into an L2-normalised embedding. The pairwise score is `cos(e_a, e_b)` (higher = more similar). |
| Odd-one-out probabilities are a `softmax` over the three "other-pair" scores |
| divided by the temperature; 2AFC compares the two reference-candidate scores. |
| The pairwise score is the model's raw output (temperature is applied at the |
| probability step). |
|
|
| | Property | Value | |
| |----------|-------| |
| | Base model | Qwen/Qwen3-VL-Embedding-8B | |
| | Pairwise score | `cos(e_a, e_b)` | |
| | Fine-tuning | LoRA (r=16, α=32) on the LLM layers | |
| | Pooling | last-token | |
| | Temperature | 0.05 (applied at the probability step) | |
| | Prompt `X` | `Represent the similarity of the image based on X.` | |
|
|
| ## Usage |
| TPIPS supports Python 3.10 and later. Install matching PyTorch and torchvision builds from the [official PyTorch installer](https://pytorch.org/get-started/locally/), then install TPIPS: |
|
|
| ```bash |
| pip install tpips |
| ``` |
|
|
| Start with the recommended `embedding` model: |
|
|
| ```python |
| import tpips |
| from PIL import Image |
| |
| model = tpips.load_model("embedding", device="cuda") |
| a = Image.open("a.jpg").convert("RGB") |
| b = Image.open("b.jpg").convert("RGB") |
| |
| similarity = model.similarity(a, b, factor="lighting") # higher is more similar |
| distance = model.distance(a, b, factor="lighting") # lower is more similar |
| ``` |
|
|
| The first call downloads the selected TPIPS checkpoint and its Qwen backbone. A CUDA GPU is recommended; FlashAttention is optional. |
|
|
|
|
| ## The TPIPS models |
|
|
| | Model | Repo | |
| |-------|------| |
| | Embedding (late fusion) | [`sywang/TPIPS-Embed-Qwen3VL-8B`](https://huggingface.co/sywang/TPIPS-Embed-Qwen3VL-8B) | |
| | Early Fusion | [`sywang/TPIPS-EarlyFusion-Qwen3VL-8B`](https://huggingface.co/sywang/TPIPS-EarlyFusion-Qwen3VL-8B) | |
| | Activation Distance | [`sywang/TPIPS-ActDiff-Qwen3VL-8B`](https://huggingface.co/sywang/TPIPS-ActDiff-Qwen3VL-8B) | |
|
|
| ## License |
|
|
| TPIPS is provided under the [Adobe Research License](https://github.com/adobe-research/TPIPS/blob/main/LICENSE) for noncommercial research use. See the license for the complete terms. |