File size: 7,893 Bytes
2b33fa9
 
72f3342
 
 
 
2b33fa9
 
 
72f3342
2b33fa9
7e15a8d
2b33fa9
 
 
 
 
72f3342
 
 
 
2b33fa9
 
 
 
3b5433c
 
 
 
2b33fa9
 
 
 
 
 
7e15a8d
2b33fa9
 
 
 
 
 
 
 
 
d2b226f
 
2b33fa9
d2b226f
2b33fa9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7e15a8d
2b33fa9
d2b226f
7e15a8d
 
 
 
 
2b33fa9
7e15a8d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2b33fa9
 
7e15a8d
 
 
 
 
2b33fa9
 
 
 
 
 
 
 
7e15a8d
 
 
 
 
 
2b33fa9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3b5433c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2b33fa9
 
 
 
 
 
 
 
 
 
 
 
 
 
a8387ae
2b33fa9
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
---
library_name: transformers
license: other
license_name: wemm-model-license
license_link: https://huggingface.co/tencent/WeMM-Embedding-9B/blob/main/LICENSE

base_model:
- Qwen/Qwen3.5-9B
pipeline_tag: feature-extraction

tags:
- sentence-transformers
- multimodal-embedding
- text-embedding
- image-embedding
- video-embedding
- mrl

language:
- zh
- en
---

# WeMM-Embedding-9B

[![Hugging Face](https://img.shields.io/badge/🤗-Hugging%20Face-yellow)](https://huggingface.co/collections/tencent/wemm-embedding)
[![Technical Report](https://img.shields.io/badge/📄-Technical%20Report-red)](https://github.com/Tencent/WeMM-Embedding/blob/main/assets/WeMM_Embedding_tech_report.pdf)
[![GitHub](https://img.shields.io/badge/GitHub-WeMM--Embedding-black?logo=github)](https://github.com/Tencent/WeMM-Embedding)

WeMM-Embedding-9B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 4,096-dimensional L2-normalized embedding. Audio input is not supported.

## Installation

```bash
pip install torch transformers==5.2.0 "qwen-vl-utils[decord]==0.0.14" \
  "sentence-transformers>=5.7.0" "accelerate>=1.1.0"
```

## Transformers

```python
import torch
from qwen_vl_utils import process_vision_info
from transformers import AutoModel, AutoProcessor

model_id = "tencent/WeMM-Embedding-9B"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModel.from_pretrained(
    model_id, trust_remote_code=True, dtype=torch.bfloat16
).cuda().eval()

messages = [{"role": "user", "content": [
    {"type": "image", "image": "/path/to/image.jpg"},
    {"type": "video", "video": "/path/to/video.mp4"},
    {"type": "text", "text": "This can be any text input."},
]}]
text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=False
)
images, videos, video_kwargs = process_vision_info(
    messages,
    image_patch_size=16,
    return_video_kwargs=True,
    return_video_metadata=True,
)
if videos is not None:
    videos, video_metadata = zip(*videos)
    videos, video_metadata = list(videos), list(video_metadata)
else:
    video_metadata = None
inputs = processor(
    text=text,
    images=images,
    videos=videos,
    video_metadata=video_metadata,
    return_tensors="pt",
    **video_kwargs,
).to("cuda")

with torch.inference_mode():
    embedding = model.embedding(**inputs)
```

Use any subset of the content items to encode text, image, or video independently.

## Sentence Transformers

```python
from sentence_transformers import SentenceTransformer

model_id = "tencent/WeMM-Embedding-9B"
model = SentenceTransformer(model_id, trust_remote_code=True)

queries = [
    "Which Llama 4 model variants are available?",
    "How is mapo tofu prepared?",
]
documents = [
    "Mapo tofu is a Sichuan dish of soft tofu simmered in a spicy, numbing sauce of chili bean paste and Sichuan peppercorn.",
    {
        "image": "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/llama4_hgf.png",
        "text": "Represent this image.",
    },
    {
        "video": "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/mapo_tofu.mp4",
        "text": "Represent this video.",
    },
]

query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# (2, 4096) (3, 4096)

similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[0.2153, 0.5843, 0.1221],
#         [0.7665, 0.2604, 0.5366]])
```

Each input is a string, a URL or path, a `PIL.Image`, or a dict combining `image`,
`video`, and `text` keys. Put `image` or `video` before `text` so the prompt matches
the ordering used above. Chat messages such as
`{"role": "user", "content": [{"type": "image", "image": ...}, {"type": "text", "text": ...}]}`
are also accepted, which is the way to interleave several images or videos in one input.

## Matryoshka Embeddings

```python
d = 256
embedding_d = torch.nn.functional.normalize(embedding[..., :d], dim=-1)
```

With Sentence Transformers, pass `truncate_dim` and let it renormalize:

```python
embeddings_d = model.encode_document(documents, truncate_dim=d, normalize_embeddings=True)
```

Use a dimension listed in `model.config.matryoshka_dimensions`.

## Serving

vLLM `0.27.0`:

```bash
MODEL_PATH=/path/to/WeMM-Embedding-9B
vllm serve "$MODEL_PATH" \
  --runner pooling \
  --chat-template "$MODEL_PATH/embedding_chat_template.jinja"
```

SGLang `0.5.9`:

```bash
MODEL_PATH=/path/to/WeMM-Embedding-9B
python patch_sglang_video.py
python -m sglang.launch_server \
  --model-path "$MODEL_PATH" \
  --is-embedding \
  --enable-precise-embedding-interpolation
```

## Evaluation

### MMEB-v2

Results on 78 datasets from Table 1 of the [technical report](https://github.com/Tencent/WeMM-Embedding/blob/main/assets/WeMM_Embedding_tech_report.pdf). Image and video tasks use Hit@1, while visual-document tasks use NDCG@5. Higher is better.

| Model | Size | AVG | Image | Video | VisDoc |
| --- | ---: | ---: | ---: | ---: | ---: |
| VLM2Vec | 2B | 47.8 | 59.7 | 29.0 | 44.0 |
| GME | 2B | 55.4 | 51.9 | 33.9 | 76.8 |
| VLM2Vec-V2 | 2B | 59.3 | 64.9 | 34.9 | 69.2 |
| Qwen3-VL-Embedding | 2B | 73.2 | 75.0 | 61.9 | 79.2 |
| DME-Small† | 2B | 74.8 | 75.9 | 65.6 | 79.9 |
| **WeMM-Embedding** | **2B** | **77.9** | **79.6** | **70.8** | **80.7** |
| **WeMM-Embedding** | **4B** | **79.2** | **80.8** | **72.1** | **82.0** |
| VLM2Vec | 8B | 53.2 | 65.5 | 34.0 | 49.1 |
| GME | 8B | 59.2 | 56.0 | 38.6 | 79.3 |
| Qwen3-VL-Embedding | 8B | 77.8 | 80.1 | 67.1 | 82.4 |
| DME-Medium† | 9B | 78.4 | 79.8 | 70.8 | 82.0 |
| **WeMM-Embedding** | **9B** | **80.6** | **81.9** | **74.3** | **83.3** |

† Closed-source leaderboard submission without publicly released model weights or a public inference endpoint.

### MMEB-v3

Results on all 190 tasks from Table 2 of the [technical report](https://github.com/Tencent/WeMM-Embedding/blob/main/assets/WeMM_Embedding_tech_report.pdf). V3-All includes the 78 MMEB-v2 tasks, 53 text tasks, 47 agent tasks, 11 audio tasks, and MCMR. Unsupported tasks are assigned a score of zero.

| Model | Size | V3-All | Text | Agent | MCMR | Audio |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| VLM2Vec-V2 | 2B | 38.3 | 24.5 | 28.7 | 4.1 | 0.0 |
| Omni-Embed-Nemotron | 3B | 43.5 | 39.2 | 36.5 | 26.1 | 36.5 |
| E5-Omni | 3B | 44.6 | 26.7 | 36.9 | 31.9 | 30.8 |
| Qwen3-VL-Embedding | 2B | 50.9 | 39.2 | 39.3 | 42.0 | 0.0 |
| **WeMM-Embedding** | **2B** | **56.0** | **45.3** | **45.1** | **42.5** | **0.0** |
| **WeMM-Embedding** | **4B** | **58.2** | **47.9** | **49.0** | **41.9** | **0.0** |
| WAVE | 7B | 26.3 | 13.7 | 11.3 | 8.9 | 31.8 |
| VLM2Vec | 8B | 32.9 | 22.2 | 19.7 | 0.9 | 0.0 |
| LCO-Embedding-Omni | 7B | 40.6 | 32.4 | 27.8 | 20.0 | 43.2 |
| GME | 8B | 43.6 | 37.1 | 35.6 | 27.3 | 0.0 |
| E5-Omni | 7B | 47.1 | 26.9 | 36.7 | 41.1 | 43.0 |
| Tianmu-Emb-Uni | 8B | 53.3 | 43.6 | 39.4 | 38.8 | 38.9 |
| Qwen3-VL-Embedding | 8B | 53.5 | 42.5 | 38.4 | 38.0 | 0.0 |
| **WeMM-Embedding** | **9B** | **59.5** | **48.8** | **51.0** | **49.3** | **0.0** |

Text results use NDCG@5; agent, MCMR, and audio results use Hit@1.

## Citation

```bibtex
@techreport{wemm_embedding_2026,
  title       = {WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report},
  author      = {{WeChat Vision}},
  institution = {Tencent Inc.},
  year        = {2026}
}
```

## License

WeMM-Embedding-9B, including the code, model parameters, and weights made publicly
available by Tencent, is licensed under the [Apache License 2.0](https://huggingface.co/tencent/WeMM-Embedding-9B/blob/main/LICENSE).
Third-party components remain subject to their respective original licenses.