OrcaRouter

GLM-5.3-Flash-MLX

An MLX build of the official GLM-5.3-Flash — 2bit-lite / 2 / 3 / 4 / 6-bit OrcaSAQ quant for Apple Silicon & MLX

Website Model Catalog Model Card License MLX Quants OrcaSAQ BF16 OrcaCode Review

One Gateway. Every Model. — Route Smarter · Ship Safer · Spend Less.

Website · Model Catalog · Model API · GitHub · OrcaCode Review · Discord · X


An MLX build of the official zai-org/GLM-5.3-Flash — a 320B-total / 18B-active MoE, the first natively multimodal model in the GLM-5 series, built on the new glm5_next hybrid architecture (sparse + linear attention with Manifold-Constrained Hyper-Connections) — quantized to MLX for Apple Silicon and the MLX CUDA backend. Five builds (2bit-lite / 2 / 3 / 4 / 6-bit) ship as subfolders, with the 4-bit build mirrored at the repo root, produced with OrcaSAQ (Sensitivity-Aware Quantization) — calibration-free and architecture-aware: sensitive tensors get more bits (shared experts +2, down_proj +1), and every tensor that was not FP8 in the base release — the 34 linear-attention layers, the sparse indexer, hyper-connections, norms, embed_tokens, lm_head, and the entire vision tower — is carried through in BF16. The four projections of the sparse-attention blocks were FP8 upstream and are quantized at the base width. Browse all models in the OrcaRouter Model Catalog; deployed as API here. Put this model to work reviewing your pull requests with OrcaCode Review.


Available quantizations

Folder Base bits Group size Size Min RAM Quality vs FP8
6-bit/ 6 64 ~296 GB 320 GB Near-lossless — best quality
4-bit/ 4 64 ~204 GB 224 GB Very good — recommended default
3-bit/ 3 32 ~184 GB 200 GB Good — aggressive but usable
2-bit/ 2 32 ~145 GB 160 GB Aggressive — best-effort
2bit-lite/ 2 128 ~102 GB 112 GB Smallest — fits 128 GB Macs & a single H200

The repository root is a mirror of 4-bit/, so orcarouter/GLM-5.3-Flash-MLX resolves to the recommended default with no subfolder needed.

Quantized from the official FP8 release. Bit-widths above are the base precision; the OrcaSAQ policy below raises sensitive tensors above the base. Group size is 64 for 4/6-bit, 32 for 2/3-bit (shared experts always use 64), and 128 throughout 2bit-lite.

No 8-bit build was produced. 2bit-lite follows a separate, memory-first recipe — see below.


OrcaSAQ — Sensitivity-Aware Quantization

OrcaSAQ is a calibration-free, architecture-aware mixed-precision quantization method designed to preserve model quality while aggressively reducing memory footprint.

Instead of running expensive calibration datasets or per-layer sensitivity searches, OrcaSAQ uses architectural and tensor-role priors to determine which weights are most sensitive to quantization. Critical components receive higher precision, while more quantization-tolerant tensors remain at the target base precision.

For MoE architectures, for example, OrcaSAQ can preserve shared experts at base +2 bits, down_proj at base +1 bit, keep gate_proj and up_proj at the base precision, and retain sensitive non-quantized components in BF16.
The result is a fast, deterministic, zero-calibration quantization pipeline that can be applied to new architectures without costly dataset-driven optimization.

Architecture-aware. Calibration-free. Quality-preserving.

Bit allocation for GLM-5.3-Flash

Component Count 2-bit 3-bit 4-bit 6-bit Policy
Expert gate_proj / up_proj 24 768 2 3 4 6 base
Expert down_proj 12 384 3 4 5 8 base +1
Dense-MLP gate_proj / up_proj (layers 0–2) 6 2 3 4 6 base
Dense-MLP down_proj (layers 0–2) 3 3 4 5 8 base +1
Shared expert gate_proj / up_proj / down_proj 129 4 5 6 8 base +2
Sparse-attn q_a_proj / q_b_proj / kv_a_proj_with_mqa / o_proj 48 2 3 4 6 base
Linear-attn layers, sparse indexer, mHC, norms, embed_tokens, lm_head, vision tower BF16 BF16 BF16 BF16 never quantized

Bits are rounded up to the nearest MLX-supported width ({2,3,4,5,6,8}). The architectural prior that selects the quantizable set is mechanical and needs no data: a tensor is re-quantized iff the FP8 release shipped it with a _scale_inv companion. For GLM-5.3-Flash that set is the MoE and dense-MLP linears plus the four projections of every deepseek_sparse_attention block — the 11 sparse layers at depth 3, 7, 11 … 43, and the MTP block, hence 12 × 4 = 48 tensors. The 34 linear_attention layers, the sparse indexer, and the vision tower were never FP8 and stay BF16.

37 338 tensors are quantized in total. The assignment is recorded in config.jsonquantization, as per-module {group_size, bits} overrides keyed by MLX module path — e.g. model.layers.3.mlp.switch_mlp.down_proj. MLX fuses a layer's routed experts into one switch_mlp, so 173 entries cover all 37 338 tensors. This is what the loader reads at load time and it is required: the top-level bits/group_size alone would mis-shape every raised-precision tensor.

The MTP (multi-token prediction) layer — layer 45 — is included inside the quantized weights rather than exported as a separate mtp/ module.


Quality vs FP8

All three tables compare each build against the full FP8 reference (dequantized to BF16 and run through the identical glm5_next forward, so the only variable is the quantization). Sizes are decimal GB, matching the file sizes in this repo.

Perplexity

Build Size PPL ΔPPL
FP8 (ref) 328 GB 2.7797
6-bit 296 GB 2.7864 +0.24 %
4-bit 204 GB 2.8620 +2.96 %
3-bit 184 GB 3.0566 +9.96 %
2-bit 145 GB 4.3622 +56.9 %
2bit-lite 102 GB 6.7018 +141 %

KL divergence & Top-1 token agreement — lower KLD and higher Top-1 = closer to the original.

Build Size Mean KLD KLD p95 Top-1 agreement
FP8 (ref) 328 GB 100 %
6-bit 296 GB 0.0063 0.0142 97.76 %
4-bit 204 GB 0.0131 0.0477 96.13 %
3-bit 184 GB 0.0421 0.1332 92.06 %
2-bit 145 GB 0.1647 0.6528 86.56 %
2bit-lite 102 GB 0.3456 1.2617 77.19 %

Weight-space fidelity (mean over sampled MoE/MLP tensors, quantized vs FP8):

Build Size Cosine sim SNR (dB) Rel. error
6-bit 296 GB 0.9998 37.3 1.6 %
4-bit 204 GB 0.9969 22.9 7.6 %
3-bit 184 GB 0.9892 17.6 14.0 %
2-bit 145 GB 0.9518 11.1 29.7 %
2bit-lite 102 GB 0.9135 7.5 42.4 %

The three measurements agree on the ordering, and they agree on where the cliff is: everything down to 3-bit degrades gently, 2-bit costs a lot, and 2bit-lite costs a lot more. Pick it for fit, not for quality.


2bit-lite/ — smallest build, for 128 GB Macs and a single H200

A memory-first build that trades quality for the smallest footprint that still loads and runs on mlx-vlm. It exists to fit machines where the regular 2-bit build (145 GB) is just too big:

  • 128 GB MacBook Pro (M4 / M5 Max) — regular 2-bit does not fit; this does, with a raised wired-memory limit.
  • A single H200 (141 GB) — regular 2-bit overflows one card; 2bit-lite fits with ~39 GB left for the KV cache.

Recipe (group_size 128 throughout):

Component Bits
Routed experts (gate_proj / up_proj / down_proj) 2
Attention MLA + dense MLP 2
Shared expert (fires on every token) 3
embed_tokens / lm_head / linear-attention path BF16

Size: ~102 GB (≈2.25 bpw on the experts). That is the floor for this architecture: MLX's affine quantizer cannot go below 2 bits, and the fused MoE experts — ~97 % of the weights — cannot use the 1-bit path.


Usage (mlx-vlm)

GLM-5.3-Flash is a vision-language model, so it runs under mlx-vlm — not mlx-lm.

pip install -U "mlx-vlm>=0.6.17"

The repo root is the 4-bit build, so the recommended default needs no subfolder — the excludes are what stop hf download from also pulling all five variant folders:

hf download orcarouter/GLM-5.3-Flash-MLX --local-dir ./GLM-5.3-Flash-MLX \
    --exclude "2bit-lite/*" "2-bit/*" "3-bit/*" "4-bit/*" "6-bit/*"

python -m mlx_vlm.generate --model ./GLM-5.3-Flash-MLX \
    --prompt "Explain quantum entanglement in one sentence." --max-tokens 256

To pick a different precision, download just that subfolder and load it from the local path:

hf download orcarouter/GLM-5.3-Flash-MLX --include "6-bit/*" --local-dir ./GLM-5.3-Flash-MLX

python -m mlx_vlm.generate --model ./GLM-5.3-Flash-MLX/6-bit \
    --prompt "Explain quantum entanglement in one sentence." --max-tokens 256

With an image or a video frame:

python -m mlx_vlm.generate --model ./GLM-5.3-Flash-MLX \
    --image photo.jpg --prompt "Describe this image." --max-tokens 256

Python API:

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template

model, processor = load("./GLM-5.3-Flash-MLX")          # repo root == 4-bit
prompt = apply_chat_template(processor, model.config, "Describe this image.", num_images=1)

print(generate(model, processor, prompt, ["photo.jpg"], max_tokens=256, verbose=True))

glm5_next runtime required. This is a brand-new architecture (hybrid sparse + linear attention, mHC, MoE). Use an mlx-vlm build that has landed glm5_next support. On Apple Silicon the Metal backend is used automatically. On the Linux CUDA backend install mlx[cuda] and set MLX_CUDA_USE_CUDNN_SDPA=0.


🐳 Run it with OrcaCode Review

OrcaRouter website Orca Code Review on GitHub

Models are only half the system.

OrcaCode Review turns every model listed on OrcaRouter into a production code-review agent:

  • reviews every PR
  • finds security + correctness issues
  • posts inline findings
  • P0/P1 can block merges
  • swap models anytime

Open model. Open harness. Open bill.


GLM-5.3-Flash

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

For the first time in the GLM series, GLM-5.3-Flash introduces a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. It also adopts Manifold-Constrained Hyper-Connections (mHC) to improve scaling efficiency, and is trained on a 30T-token multimodal corpus.

Model Overview

  • Type: Mixture-of-Experts multimodal Causal LM (glm5_next) with vision encoder
  • Parameters: 320B total · 18B active per token
  • Experts: 288 routed + 1 shared, top-8 routing
  • Layers: 45 (3 dense + 42 MoE) · 11 deepseek_sparse_attention + 34 linear_attention, plus 1 MTP layer (also MoE, with its own sparse attention)
  • Attention: hybrid sparse + linear attention with a learned sparse indexer, plus Manifold-Constrained Hyper-Connections (mHC)
  • Precision (base release): FP8 (block-wise e4m3, 128×128), dynamic activation scheme
  • Modalities: text, image, video

Best Practices

  • Sampling: follow the official GLM-5.3-Flash guidance; long-horizon agentic and coding tasks benefit from generous max-tokens headroom.
  • Long context: the hybrid sparse/linear attention is designed for low-cost long-context serving; give the runtime enough KV budget for your target length.
  • Pick a precision: 6-bit for near-lossless, 4-bit (repo root) as the everyday default, 3-bit and 2-bit when memory is the binding constraint, 2bit-lite only when nothing else fits (see the quality tables above — the drop from 2-bit to 2bit-lite is steep).

Citation

@misc{glm5team2026glm5vibecodingagentic,
      title={GLM-5: from Vibe Coding to Agentic Engineering},
      author={GLM-5-Team and Aohan Zeng and Xin Lv and Zhenyu Hou and Zhengxiao Du and others},
      year={2026},
      eprint={2602.15763},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2602.15763},
}

License

MIT, inherited from zai-org/GLM-5.3-Flash. Quantization does not change the underlying license obligations.

Downloads last month
-
Safetensors
Model size
59B params
Tensor type
U32
·
BF16
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for orcarouter/GLM-5.3-Flash-MLX

Quantized
(24)
this model

Paper for orcarouter/GLM-5.3-Flash-MLX