AX Qwen3 Coder Next MLX 4-bit

Parameter count: approximately 79.67B logical parameters (80B total, 3B active per token). 4-bit is the quantization precision, not a 4B model-size claim.

This is a revision-pinned, transparent mirror of mlx-community/Qwen3-Coder-Next-4bit at commit 7b9321eabb85ce79625cac3f61ea691e4ea984b5.

The model weights, index, configuration, tokenizer, chat template, generation configuration, and tool-parser files are byte-identical to that upstream revision. AutomatosX did not fine-tune, merge, re-quantize, or otherwise alter the model artifacts. We add only mirror documentation, a copy of the declared Apache 2.0 license, and machine-readable provenance.

Model details

  • Base model: Qwen/Qwen3-Coder-Next
  • Format: MLX Safetensors for Apple Silicon
  • Architecture: Qwen3NextForCausalLM, mixture of experts
  • Main quantization: 4-bit affine, group size 64
  • Quantization exceptions: router and shared-expert gate tensors remain 8-bit, as defined by the unchanged upstream config
  • Layers: 48
  • Experts: 512 total, 10 selected per token
  • Configured context limit: 262,144 tokens
  • Weight shards: 9, totaling 44,844,286,500 bytes
  • Upstream conversion tool: mlx-lm 0.30.5

Download

hf download AutomatosX/AX-Qwen3-Coder-Next-MLX-4bit \
  --local-dir ./AX-Qwen3-Coder-Next-MLX-4bit

Use with MLX-LM

pip install mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3-Coder-Next-MLX-4bit \
  --prompt "Write a Python function that merges two sorted lists."

Applications should apply the included chat template for conversational or tool-using prompts.

Serve with AX Engine

You can also serve the downloaded model through the OpenAI-compatible API in AX Engine:

ax-engine serve ./AX-Qwen3-Coder-Next-MLX-4bit --port 31418

Mirror policy and provenance

This release is meant to group a required upstream artifact under the AutomatosX catalog, not to claim a new conversion. UPSTREAM_README.md preserves the original upstream model card. ax_provenance.json pins the source commit and records SHA-256 values and sizes for every mirrored artifact.

The local Hugging Face cache contained an AX-generated model-manifest.json; it is not present in the pinned upstream repository and was deliberately not published here.

This is a standard direct-decoding model. It does not contain an MTP head and no MTP conversion was applied.

License

The upstream model metadata declares Apache License 2.0. See LICENSE, the upstream model card, and the base-model card for limitations and responsible-use guidance.

Downloads last month
186
Safetensors
Model size
80B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AutomatosX/AX-Qwen3-Coder-Next-MLX-4bit

Quantized
(111)
this model

Collection including AutomatosX/AX-Qwen3-Coder-Next-MLX-4bit