GLM-5.3-Flash-AWQ-W4A16

Quantized version of zai-org/GLM-5.3-Flash.

If you come here with older cards like A100, A6000 or 3090, consider using our vLLM fork which has served billions of tokens for frontier models with older GPUs.

What is quantized

Weight-only INT4 (symmetric, group size 128) via AWQ, stored in the compressed-tensors pack-quantized format. Calibrated from the official FP8 release (dequantized to BF16 first).

Quantized (INT4 W4A16):

  • Routed MoE experts in the 42 MoE decoder layers (layers 3-44): mlp.experts.{0..287}.{gate_proj, up_proj, down_proj} (~312B of the 321B parameters)

Kept in BF16 (not quantized):

  • Token embedding (embed_tokens) and lm_head
  • Linear attention / KDA (self_attn.{q,k,v,o}_proj, forget gate) and DSA/MLA attention (q_a/q_b/kv_a/kv_b_proj, indexer)
  • Hyper-connections (attn_hc.*, ffn_hc.*)
  • MoE router (mlp.gate) and shared experts (mlp.shared_experts.*)
  • Dense MLPs of the first 3 layers
  • Vision encoder (model.visual.*)
  • NextN/MTP layer (layers.45.*, dequantized from the FP8 source to BF16)
Downloads last month
126
Safetensors
Model size
321B params
Tensor type
F32
I32
BF16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for wtdcode/GLM-5.3-Flash-AWQ-W4A16

Quantized
(50)
this model