GLM-5.2-MLX-nvfp4

Runtime — updated 2026-08-28: load with --trust-remote-code

This repository now bundles glm_moe_dsa.py (declared via model_file in config.json), a fixed runtime for this architecture, and needs it:

mlx_lm.generate --model pipenetwork/GLM-5.2-MLX-nvfp4 --trust-remote-code --prompt "..." --max-tokens 300

mlx-lm's own glm_moe_dsa builds a lightning indexer on all 78 layers, but GLM-5.2 ships indexer weights on 21 (indexer_types: the other 57 "shared" layers reuse the previous full layer's top-k selection). mlx_lm.load loads leniently and left those 57 indexers at random initialisation. Prompts up to 2048 tokens were unaffected (the indexer is bypassed below index_topk); beyond that, 57 of 78 layers attended to keys chosen by random projections. The bundled runtime implements the schedule as the reference does (plus fp32 indexer scores and router logits and the indexer LayerNorm epsilon); tiny-config parity against transformers 5.16 is 4e-7 with the sparse path live, and a strict load of this checkpoint reports zero missing and zero unexpected tensors. Details, tests and the GLM-5.3 builds made with it: github.com/PipeNetwork/glm53-mlx. The weights are unchanged.

An MLX conversion of zai-org/GLM-5.2 quantized to NVFP4 (4-bit FP4, group size 16) for Apple Silicon with mlx-lm.

This is the MLX analog of NVIDIA's nvidia/GLM-5.2-NVFP4. NVIDIA's checkpoint stores weights in ModelOpt-packed NVFP4 that mlx-lm cannot read directly, so this build was produced by quantizing the bf16 base with MLX's own NVFP4 mode (--q-mode nvfp4 --q-group-size 16).

  • Base model: zai-org/GLM-5.2 (GlmMoeDsaForCausalLM, 753B total / ~40B active MoE, text-only)
  • Format: MLX, NVFP4 (4-bit FP4, group size 16)
  • Approx. size on disk: 390G
  • Converted with: mlx-lm 0.31.2

Usage

pip install -U mlx-lm
mlx_lm.generate --model pipenetwork/GLM-5.2-MLX-nvfp4 --prompt "Explain mixture-of-experts in one sentence." --max-tokens 128

License

MIT, inherited from the base model.

Downloads last month
686
Safetensors
Model size
743B params
Tensor type
U32
·
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pipenetwork/GLM-5.2-MLX-nvfp4

Base model

zai-org/GLM-5.2
Quantized
(146)
this model

Collection including pipenetwork/GLM-5.2-MLX-nvfp4