File size: 5,750 Bytes
72a1733
 
 
 
 
 
 
 
 
 
 
 
 
daec459
72a1733
 
 
b10396b
72a1733
 
 
 
 
 
 
deaa046
72a1733
 
 
 
daec459
 
 
72a1733
daec459
72a1733
daec459
 
 
72a1733
daec459
 
72a1733
 
 
072b5e1
ac12dbb
72a1733
 
072b5e1
 
 
72a1733
 
 
 
 
 
 
 
072b5e1
daec459
 
 
72a1733
daec459
72a1733
 
daec459
 
72a1733
daec459
57b4dcf
72a1733
daec459
72a1733
 
 
 
 
072b5e1
72a1733
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
072b5e1
72a1733
 
 
 
072b5e1
 
 
 
 
 
72a1733
 
072b5e1
 
 
 
 
daec459
072b5e1
 
 
72a1733
 
072b5e1
 
 
 
20fe80b
09711ee
 
072b5e1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
09711ee
 
072b5e1
20fe80b
072b5e1
72a1733
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
---
license: mit
library_name: sglang
pipeline_tag: image-text-to-text
tags:
  - multimodal
  - vision-language
  - glm
  - sglang
base_model:
  - zai-org/GLM-5.2
---

# GLM-5.2-Vision (FP8)

**GLM-5.2 with sight.** A vision-language model that bolts the MoonViT vision encoder from
[Kimi-K2.6](https://huggingface.co/moonshotai/Kimi-K2.6) onto
[GLM-5.2](https://huggingface.co/zai-org/GLM-5.2) through a trained PatchMerger projector.

GLM-5.2 is a strong open reasoning model with no vision input. This checkpoint adds it,
without touching a single GLM weight: the text backbone and the vision tower are both frozen
and byte-identical to their upstream releases. The only newly-trained parameters are the
**49.5M-parameter projector** that maps MoonViT's 1152-dim patch embeddings into GLM's 6144-dim
token space.

| Component | Detail |
|---|---|
| Text backbone | GLM-5.2 (744B total / A40B active, MoE + MLA + DSA sparse attention) — **frozen** |
| Vision tower | MoonViT-3d from Kimi-K2.6, 27 layers, 1152-dim — **frozen** |
| Projector | PatchMerger MLP (`pre_norm → linear_1 → GELU → linear_2`), 1152→4608→6144 — **trained** |
| Text weights | block-FP8, from [`zai-org/GLM-5.2-FP8`](https://huggingface.co/zai-org/GLM-5.2-FP8) |
| Size | ~757 GB |
| Hardware | 8×B200 or 8×H200 |
| Image tokens | up to 4096 per image (16384 MoonViT patches, 2×2 merge) |
| Max context | 1048576 (1M tokens) |

The vision tower and projector are **bf16** — only the GLM text Linears are
quantized. This is the most thoroughly validated build; if you are unsure which
variant to use, use this one.

At ~757 GB the weights do not fit on four GPUs. For a 4-GPU deployment use
[`baseten/GLM-5.2-Vision-NVFP4`](https://huggingface.co/baseten/GLM-5.2-Vision-NVFP4).

## Quickstart

SGLang needs a small out-of-tree plugin because `Glm5vForConditionalGeneration` is not yet an
upstream architecture. It ships inside this repo, so there is nothing else to clone:

```bash
uvx --from huggingface-hub hf download baseten/GLM-5.2-Vision-FP8 \
  --include 'plugins/*' --local-dir ./glm5v
uv pip install ./glm5v/plugins
```

### SGLang

```bash
export SGLANG_EXTERNAL_MODEL_PACKAGE=sglang_glm5v
export SGLANG_EXTERNAL_MM_PROCESSOR_PACKAGE=sglang_glm5v
export SGLANG_EXTERNAL_MM_MODEL_ARCH=Glm5vForConditionalGeneration
python -m sglang_glm5v.patch
```

#### 8×B200 / 8×H200 — full 1M context

```bash
python -m sglang.launch_server \
  --model-path baseten/GLM-5.2-Vision-FP8 --trust-remote-code \
  --tp-size 8 \
  --attention-backend dsa --mm-attention-backend sdpa \
  --kv-cache-dtype fp8_e4m3 --page-size 64 \
  --mem-fraction-static 0.85 \
  --context-length 1048576 \
  --reasoning-parser glm45 --tool-call-parser glm47 \
  --served-model-name glm-5.2-vision \
  --port 30000
```

### Query it

Standard OpenAI multimodal messages deliver the image as `image_url`:

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="none")
r = client.chat.completions.create(
    model="glm-5.2-vision",
    messages=[{"role": "user", "content": [
        {"type": "image_url", "image_url": {"url": "https://ultralytics.com/images/bus.jpg"}},
        {"type": "text", "text": "Describe this image in detail."},
    ]}],
    temperature=1.0, top_p=0.95, max_tokens=512,
)
print(r.choices[0].message.content)
```

GLM-5.2 is a reasoning model: with `--reasoning-parser glm45`, the chain of thought arrives in
`message.reasoning_content` and the answer in `message.content`.

## Deploy on Baseten

The repository includes ready-to-push [Truss](https://truss.baseten.co) configs. The only
credential you need is an API key for your own Baseten account; no Hugging Face token or
pre-created Baseten secret is required.

1. Install [`uv`](https://docs.astral.sh/uv/getting-started/installation/) and create a Baseten API key.
2. Export the key, download the small Truss directory, and deploy the model:

```bash
export BASETEN_API_KEY="your-baseten-api-key"
uvx truss login --api-key "$BASETEN_API_KEY" --remote baseten --non-interactive

uvx --from huggingface-hub hf download baseten/GLM-5.2-Vision-FP8 \
  --include 'truss/*' --local-dir ./glm5v
cd glm5v/truss

# FP8 requires 8×B200 and provides the full 1M-token context.
uvx truss push --remote baseten --config config.yaml --wait --output json
```

The command creates a new model and published deployment in your Baseten account and prints
JSON containing `model_id`, `model_version_id`, `predict_url`, and `logs_url`. It does not
promote the deployment to production.

Set `PREDICT_URL` to the returned `predict_url`, then query the model:

```bash
export PREDICT_URL="https://model-...api.baseten.co/deployment/.../predict"

curl -fsS "$PREDICT_URL" \
  -H "Authorization: Api-Key $BASETEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.2-vision",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "image_url", "image_url": {"url": "https://ultralytics.com/images/bus.jpg"}},
        {"type": "text", "text": "Describe this image in detail."}
      ]
    }],
    "max_tokens": 512,
    "temperature": 1.0,
    "top_p": 0.95
  }'
```

The first deployment downloads about 757 GB of weights and initializes SGLang, so startup can
take several minutes.

## License

MIT, following both parents: GLM-5.2 (MIT) and Kimi-K2.6 (Modified MIT). The projector weights
are released under MIT. Redistributed upstream weights remain under their original terms.

## Acknowledgements

Built on [Z.ai](https://huggingface.co/zai-org)'s GLM-5.2 and
[Moonshot AI](https://huggingface.co/moonshotai)'s Kimi-K2.6. Neither team was involved in this
work; please do not direct issues with this checkpoint to them.