GLM-5.3-Flash-NVFP4 / README.md
jon1012's picture
add README.md
462753f verified
|
Raw
History Blame Contribute Delete
6.15 kB
---
license: mit
base_model: zai-org/GLM-5.3-Flash
base_model_relation: quantized
quantized_by: LibertAIDAI
tags:
- nvfp4
- blackwell
- vllm
- glm
- glm-5
- glm5_next
- moe
- multimodal
- modelopt
language:
- en
- zh
pipeline_tag: image-text-to-text
---
# GLM-5.3-Flash NVFP4 (weight-only)
NVFP4 quantization of [zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash) — Z.ai's **320B / 18B-active** natively-multimodal MoE, the first GLM to combine **sparse and linear attention** — for serving on [vLLM](https://github.com/vllm-project/vllm).
This is a **weight-only NVFP4 (NVFP4-A16)** checkpoint. The **routed-expert FFN** tensors — 97% of the model's parameters — are quantized to **NVFP4** (NVIDIA's 4-bit float: E2M1 with FP8-E4M3 per-16-block scales and an FP32 per-tensor global scale). Every outlier-sensitive path is deliberately kept in **BF16**: both attention flavours (the 34 KDA linear-attention layers *and* the 11 DeepSeek-sparse-attention layers, including the sparse indexer), the **entire vision tower**, the shared experts, the MoE routers, the dense-MLP layers, the MTP head, the Manifold-Constrained Hyper-Connection tensors, token embeddings, `lm_head`, and all norms. Activations stay BF16 (no activation quantization).
Because the experts dominate the footprint, quantizing only them takes the model from **598.5 GiB → ~181 GiB (−70%)** while protecting quality where it matters most — and leaving the vision tower untouched keeps multimodal behaviour bit-identical to the source.
Produced with [NVIDIA ModelOpt](https://github.com/NVIDIA/TensorRT-Model-Optimizer) 0.45.0 via a memory-frugal shard-streaming pass (no calibration data — weight-only NVFP4 derives its scales from the weights themselves), so the packing is exactly what vLLM's NVFP4 loader expects.
> ### ⚠️ Status: use the dedicated vLLM image, not a release build
> GLM-5.3-Flash was released on **2026-08-26** and its `glm5_next` architecture is **not yet in vLLM `main`** — support is in-flight as [vllm#53906](https://github.com/vllm-project/vllm/pull/53906) (and [sglang#36507](https://github.com/sgl-project/sglang/pull/36507)). Until it lands, vLLM ships support in **per-model docker images**:
>
> | | image |
> |---|---|
> | x86_64, CUDA 13 | `vllm/vllm-openai:glm53-flash-x86_64-cu130` |
> | **arm64 (GH200/GB200/GB10), CUDA 13** | `vllm/vllm-openai:glm53-flash-arm64-cu130` |
> | multi-arch | `vllm/vllm-openai:glm53-flash` |
>
> The vendor recipe lists **H100 / B200 / GB200 as verified**; sm_121 (GB10) is *not* on that list, so expect the usual FP4-MoE kernel-gap dance there. **llama.cpp has no `glm5_next` support**, so there is no GGUF yet.
## About LibertAI
[LibertAI](https://libertai.io) is a decentralized AI platform — private inference, an OpenAI-compatible API, and a chat UI, all running on community GPUs over [Aleph Cloud](https://aleph.cloud) instead of a single company's servers. No accounts required to chat, no logs sent home, and the same models you'd self-host are available behind a sovereign endpoint.
If you want to put this model (or any other) to work as an autonomous agent without running your own infrastructure, check out [**LiberClaw**](https://liberclaw.ai) — Hermes-style agents hosted on Aleph Cloud with LibertAI inference. Free tier: 2 agents, no credit card, 5 minutes to deploy. Open source.
> **Why NVFP4?** On NVIDIA Blackwell GPUs (RTX 50-series, B100/B200, GB10/DGX Spark) NVFP4 weights run on native FP4 tensor-core kernels, and the 4-bit expert weights cut both VRAM and the memory bandwidth that dominates MoE decode. This gives format parity with the NVFP4 path used by vLLM / SGLang / TensorRT-LLM.
## What's quantized
| Tensor group | Precision | Count |
|---|---|---|
| Routed-expert FFN (`language_model.layers.*.mlp.experts.*.{gate,up,down}_proj`) | **NVFP4** (g16, weight-only) | 37,152 |
| KDA linear attention (34 layers), DeepSeek sparse attention + indexer (11 layers) | BF16 | — |
| Vision tower (`model.visual.*`), shared experts, routers, dense/MTP MLP, mHC, embeddings, `lm_head`, norms | BF16 | 1,618 |
**38,770 tensors total.** Every non-expert Linear is listed in `config.json`'s `ignore`, so vLLM keeps it in BF16.
## Size
| | BF16 source | This checkpoint |
|---|---|---|
| Routed experts (311.65B params) | 623 GB | **175 GB** (NVFP4) |
| Everything else (9.67B params) | 19 GB | 19 GB (BF16) |
| **Total** | **598.5 GiB** | **≈181 GiB** |
That is the difference between "needs a GB200 tray" and "fits on two GB10 desktops."
## Usage (vLLM)
Run the dedicated image (see the status box above for the tag matching your arch):
```bash
docker run --gpus all --ipc=host -p 8000:8000 \
vllm/vllm-openai:glm53-flash-x86_64-cu130 \
--model LibertAIDAI/GLM-5.3-Flash-NVFP4 \
--tensor-parallel-size 4 \
--tool-call-parser glm47 --enable-auto-tool-choice \
--reasoning-parser glm45
```
The checkpoint keeps its **MTP layer** in BF16, so speculative decoding works as in the vendor recipe:
```bash
--speculative-config '{"method":"mtp","num_speculative_tokens":5}'
```
Engine init on a 320B MoE is slow — the recipe sets `VLLM_ENGINE_READY_TIMEOUT_S=3600`; do the same or the server will be killed mid-warmup.
On **GB10 / DGX Spark (sm_121)**, use the `arm64-cu130` tag. If you hit `cudaErrorNoKernelImageForDevice` from the native FP4 MoE kernels, fall back to the **marlin** MoE backend (dequant-to-FP16, known-good on sm_121):
```bash
--moe-backend marlin --enforce-eager
```
## Provenance & method
- Base: `zai-org/GLM-5.3-Flash-BF16` (BF16, 120 shards, 598.5 GiB).
- Quantizer: shard-streaming ModelOpt `NVFP4QTensor` weight-only pass, CPU-only, never more than one shard resident. Expert FFN → NVFP4, everything else copied through in BF16.
- Partition check: every Linear weight is provably quantized **XOR** ignored — no tensor both, none uncovered.
- Verification: per-expert round-trip **cosine ≈ 0.9967, relative error ≈ 0.0925** vs the BF16 source.
_Quantized by [LibertAI](https://libertai.io). Not affiliated with Z.ai / Zhipu. MIT licensed, as is the base model._