GLM-5.3-Flash-AWQ-W4A16
Quantized version of zai-org/GLM-5.3-Flash.
If you come here with older cards like A100, A6000 or 3090, consider using our vLLM fork which has served billions of tokens for frontier models with older GPUs.
What is quantized
Weight-only INT4 (symmetric, group size 128) via AWQ, stored in the compressed-tensors pack-quantized format. Calibrated from the official FP8 release (dequantized to BF16 first).
Quantized (INT4 W4A16):
- Routed MoE experts in the 42 MoE decoder layers (layers 3-44):
mlp.experts.{0..287}.{gate_proj, up_proj, down_proj}(~312B of the 321B parameters)
Kept in BF16 (not quantized):
- Token embedding (
embed_tokens) andlm_head - Linear attention / KDA (
self_attn.{q,k,v,o}_proj, forget gate) and DSA/MLA attention (q_a/q_b/kv_a/kv_b_proj, indexer) - Hyper-connections (
attn_hc.*,ffn_hc.*) - MoE router (
mlp.gate) and shared experts (mlp.shared_experts.*) - Dense MLPs of the first 3 layers
- Vision encoder (
model.visual.*) - NextN/MTP layer (
layers.45.*, dequantized from the FP8 source to BF16)
- Downloads last month
- 126
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support
Model tree for wtdcode/GLM-5.3-Flash-AWQ-W4A16
Base model
zai-org/GLM-5.3-Flash