GLM-5.3-Flash - 50% Expert Pruned
50% of the MoE experts pruned by router-weight magnitude using compressed-tensors.
- Base model: zai-org/GLM-5.3-Flash
- Sparsity: 50% of routed experts removed per layer
- Layers pruned: all 43 MoE layers (body layers 3-44)
- MTP layer: retained exactly as-is
- Shared experts: untouched
- Vision tower: untouched (dense ViT MLP, no MoE experts)
Reproduction
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support
Model tree for inference-optimization/GLM-5.3-Flash-MEP50
Base model
zai-org/GLM-5.3-Flash