GLM-5.3 MLX
Apple Silicon (MLX) builds of GLM-5.3 (744B) with a fixed glm_moe_dsa runtime: github.com/PipeNetwork/glm53-mlx
Text Generation • 119B • Updated • 80 • 1Note 427.7 GB, 512 GB Mac. ppl 2.7420 — the build to use: 4.3% better than uniform 4-bit for 9 GB more.
pipenetwork/GLM-5.3-MLX-4bit
Text Generation • 116B • Updated • 47 • 1Note 418.6 GB, 512 GB Mac (tight). ppl 2.8636.
pipenetwork/GLM-5.3-MLX-mixed-3_6bit
Text Generation • 95B • Updated • 71Note 332.6 GB, 384 GB-class. ppl 3.0338 (+5.9% vs 4-bit): 3-bit experts compound over 78 layers.
pipenetwork/GLM-5.3-MLX-8bit
Text Generation • 209B • Updated • 3Note ~800 GB, two machines. Closest to bf16 on the per-layer ladder (free-running error 0.131, cos 0.989).
pipenetwork/GLM-5.3-MLX-6bit
Text Generation • 163B • Updated • 14Note 604 GB, two machines. Ladder 0.167 — better than the upstream FP8 release (0.173).
pipenetwork/GLM-5.3-MLX-5bit
Text Generation • 140B • Updated • 80Note 512 GB, two machines. Ladder 0.225.
pipenetwork/GLM-5.3-REAP25-MLX-4bit
Text Generation • 88B • Updated • 7Note 316.6 GB. 192 of 256 experts per layer. ppl 3.2872 (+14.8% vs 4-bit).
pipenetwork/GLM-5.3-REAP37-MLX-4bit
Text Generation • 74B • Updated • 7Note 267.2 GB. 161 of 256 experts. ppl 3.8517 (+34.5% vs 4-bit).
pipenetwork/GLM-5.3-REAP50-MLX-4bit
Text Generation • 60B • UpdatedNote 214.6 GB, 256 GB Mac. 128 of 256 experts. ppl 5.0295 (+75.6% vs 4-bit) — coherent, but GLM-5.3 tolerates pruning poorly; the 512 GB builds are a different class.