Feature Extraction
Transformers
Safetensors
sentence-transformers
Chinese
English
qwen3_5
image-text-to-text
multimodal-embedding
text-embedding
image-embedding
video-embedding
mrl
custom_code
Instructions to use tencent/WeMM-Embedding-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use tencent/WeMM-Embedding-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="tencent/WeMM-Embedding-9B", trust_remote_code=True)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("tencent/WeMM-Embedding-9B", trust_remote_code=True) model = AutoModelForMultimodalLM.from_pretrained("tencent/WeMM-Embedding-9B", trust_remote_code=True, device_map="auto") - sentence-transformers
How to use tencent/WeMM-Embedding-9B with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("tencent/WeMM-Embedding-9B", trust_remote_code=True) sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
File size: 7,893 Bytes
2b33fa9 72f3342 2b33fa9 72f3342 2b33fa9 7e15a8d 2b33fa9 72f3342 2b33fa9 3b5433c 2b33fa9 7e15a8d 2b33fa9 d2b226f 2b33fa9 d2b226f 2b33fa9 7e15a8d 2b33fa9 d2b226f 7e15a8d 2b33fa9 7e15a8d 2b33fa9 7e15a8d 2b33fa9 7e15a8d 2b33fa9 3b5433c 2b33fa9 a8387ae 2b33fa9 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 | ---
library_name: transformers
license: other
license_name: wemm-model-license
license_link: https://huggingface.co/tencent/WeMM-Embedding-9B/blob/main/LICENSE
base_model:
- Qwen/Qwen3.5-9B
pipeline_tag: feature-extraction
tags:
- sentence-transformers
- multimodal-embedding
- text-embedding
- image-embedding
- video-embedding
- mrl
language:
- zh
- en
---
# WeMM-Embedding-9B
[](https://huggingface.co/collections/tencent/wemm-embedding)
[](https://github.com/Tencent/WeMM-Embedding/blob/main/assets/WeMM_Embedding_tech_report.pdf)
[](https://github.com/Tencent/WeMM-Embedding)
WeMM-Embedding-9B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 4,096-dimensional L2-normalized embedding. Audio input is not supported.
## Installation
```bash
pip install torch transformers==5.2.0 "qwen-vl-utils[decord]==0.0.14" \
"sentence-transformers>=5.7.0" "accelerate>=1.1.0"
```
## Transformers
```python
import torch
from qwen_vl_utils import process_vision_info
from transformers import AutoModel, AutoProcessor
model_id = "tencent/WeMM-Embedding-9B"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModel.from_pretrained(
model_id, trust_remote_code=True, dtype=torch.bfloat16
).cuda().eval()
messages = [{"role": "user", "content": [
{"type": "image", "image": "/path/to/image.jpg"},
{"type": "video", "video": "/path/to/video.mp4"},
{"type": "text", "text": "This can be any text input."},
]}]
text = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=False
)
images, videos, video_kwargs = process_vision_info(
messages,
image_patch_size=16,
return_video_kwargs=True,
return_video_metadata=True,
)
if videos is not None:
videos, video_metadata = zip(*videos)
videos, video_metadata = list(videos), list(video_metadata)
else:
video_metadata = None
inputs = processor(
text=text,
images=images,
videos=videos,
video_metadata=video_metadata,
return_tensors="pt",
**video_kwargs,
).to("cuda")
with torch.inference_mode():
embedding = model.embedding(**inputs)
```
Use any subset of the content items to encode text, image, or video independently.
## Sentence Transformers
```python
from sentence_transformers import SentenceTransformer
model_id = "tencent/WeMM-Embedding-9B"
model = SentenceTransformer(model_id, trust_remote_code=True)
queries = [
"Which Llama 4 model variants are available?",
"How is mapo tofu prepared?",
]
documents = [
"Mapo tofu is a Sichuan dish of soft tofu simmered in a spicy, numbing sauce of chili bean paste and Sichuan peppercorn.",
{
"image": "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/llama4_hgf.png",
"text": "Represent this image.",
},
{
"video": "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/mapo_tofu.mp4",
"text": "Represent this video.",
},
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# (2, 4096) (3, 4096)
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[0.2153, 0.5843, 0.1221],
# [0.7665, 0.2604, 0.5366]])
```
Each input is a string, a URL or path, a `PIL.Image`, or a dict combining `image`,
`video`, and `text` keys. Put `image` or `video` before `text` so the prompt matches
the ordering used above. Chat messages such as
`{"role": "user", "content": [{"type": "image", "image": ...}, {"type": "text", "text": ...}]}`
are also accepted, which is the way to interleave several images or videos in one input.
## Matryoshka Embeddings
```python
d = 256
embedding_d = torch.nn.functional.normalize(embedding[..., :d], dim=-1)
```
With Sentence Transformers, pass `truncate_dim` and let it renormalize:
```python
embeddings_d = model.encode_document(documents, truncate_dim=d, normalize_embeddings=True)
```
Use a dimension listed in `model.config.matryoshka_dimensions`.
## Serving
vLLM `0.27.0`:
```bash
MODEL_PATH=/path/to/WeMM-Embedding-9B
vllm serve "$MODEL_PATH" \
--runner pooling \
--chat-template "$MODEL_PATH/embedding_chat_template.jinja"
```
SGLang `0.5.9`:
```bash
MODEL_PATH=/path/to/WeMM-Embedding-9B
python patch_sglang_video.py
python -m sglang.launch_server \
--model-path "$MODEL_PATH" \
--is-embedding \
--enable-precise-embedding-interpolation
```
## Evaluation
### MMEB-v2
Results on 78 datasets from Table 1 of the [technical report](https://github.com/Tencent/WeMM-Embedding/blob/main/assets/WeMM_Embedding_tech_report.pdf). Image and video tasks use Hit@1, while visual-document tasks use NDCG@5. Higher is better.
| Model | Size | AVG | Image | Video | VisDoc |
| --- | ---: | ---: | ---: | ---: | ---: |
| VLM2Vec | 2B | 47.8 | 59.7 | 29.0 | 44.0 |
| GME | 2B | 55.4 | 51.9 | 33.9 | 76.8 |
| VLM2Vec-V2 | 2B | 59.3 | 64.9 | 34.9 | 69.2 |
| Qwen3-VL-Embedding | 2B | 73.2 | 75.0 | 61.9 | 79.2 |
| DME-Small† | 2B | 74.8 | 75.9 | 65.6 | 79.9 |
| **WeMM-Embedding** | **2B** | **77.9** | **79.6** | **70.8** | **80.7** |
| **WeMM-Embedding** | **4B** | **79.2** | **80.8** | **72.1** | **82.0** |
| VLM2Vec | 8B | 53.2 | 65.5 | 34.0 | 49.1 |
| GME | 8B | 59.2 | 56.0 | 38.6 | 79.3 |
| Qwen3-VL-Embedding | 8B | 77.8 | 80.1 | 67.1 | 82.4 |
| DME-Medium† | 9B | 78.4 | 79.8 | 70.8 | 82.0 |
| **WeMM-Embedding** | **9B** | **80.6** | **81.9** | **74.3** | **83.3** |
† Closed-source leaderboard submission without publicly released model weights or a public inference endpoint.
### MMEB-v3
Results on all 190 tasks from Table 2 of the [technical report](https://github.com/Tencent/WeMM-Embedding/blob/main/assets/WeMM_Embedding_tech_report.pdf). V3-All includes the 78 MMEB-v2 tasks, 53 text tasks, 47 agent tasks, 11 audio tasks, and MCMR. Unsupported tasks are assigned a score of zero.
| Model | Size | V3-All | Text | Agent | MCMR | Audio |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| VLM2Vec-V2 | 2B | 38.3 | 24.5 | 28.7 | 4.1 | 0.0 |
| Omni-Embed-Nemotron | 3B | 43.5 | 39.2 | 36.5 | 26.1 | 36.5 |
| E5-Omni | 3B | 44.6 | 26.7 | 36.9 | 31.9 | 30.8 |
| Qwen3-VL-Embedding | 2B | 50.9 | 39.2 | 39.3 | 42.0 | 0.0 |
| **WeMM-Embedding** | **2B** | **56.0** | **45.3** | **45.1** | **42.5** | **0.0** |
| **WeMM-Embedding** | **4B** | **58.2** | **47.9** | **49.0** | **41.9** | **0.0** |
| WAVE | 7B | 26.3 | 13.7 | 11.3 | 8.9 | 31.8 |
| VLM2Vec | 8B | 32.9 | 22.2 | 19.7 | 0.9 | 0.0 |
| LCO-Embedding-Omni | 7B | 40.6 | 32.4 | 27.8 | 20.0 | 43.2 |
| GME | 8B | 43.6 | 37.1 | 35.6 | 27.3 | 0.0 |
| E5-Omni | 7B | 47.1 | 26.9 | 36.7 | 41.1 | 43.0 |
| Tianmu-Emb-Uni | 8B | 53.3 | 43.6 | 39.4 | 38.8 | 38.9 |
| Qwen3-VL-Embedding | 8B | 53.5 | 42.5 | 38.4 | 38.0 | 0.0 |
| **WeMM-Embedding** | **9B** | **59.5** | **48.8** | **51.0** | **49.3** | **0.0** |
Text results use NDCG@5; agent, MCMR, and audio results use Hit@1.
## Citation
```bibtex
@techreport{wemm_embedding_2026,
title = {WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report},
author = {{WeChat Vision}},
institution = {Tencent Inc.},
year = {2026}
}
```
## License
WeMM-Embedding-9B, including the code, model parameters, and weights made publicly
available by Tencent, is licensed under the [Apache License 2.0](https://huggingface.co/tencent/WeMM-Embedding-9B/blob/main/LICENSE).
Third-party components remain subject to their respective original licenses.
|