AutomatosX's picture
Add files using upload-large-folder tool
bfe8589 verified
|
Raw
History Blame Contribute Delete
3.71 kB
---
license: apache-2.0
base_model:
- Qwen/Qwen3-Coder-Next
- mlx-community/Qwen3-Coder-Next-bf16
pipeline_tag: text-generation
library_name: mlx
tags:
- mlx
- apple-silicon
- qwen3-next
- coding
- mixture-of-experts
- quantized
- mixed-precision
- optiq
- 4-bit
- 8-bit
- automatosx
- ax-engine
---
# AX Qwen3 Coder Next MLX OptiQ 4-bit
> **Parameter count:** 79,674,391,296 logical parameters (approximately
> 79.67B total, ~3B active per token). `4-bit` names the low OptiQ
> tier, not a 4B model-size claim.
This is an MLX OptiQ mixed-precision release of Qwen3-Coder-Next for Apple
Silicon. AutomatosX converted it from the revision-pinned BF16 MLX source
[mlx-community/Qwen3-Coder-Next-bf16](https://huggingface.co/mlx-community/Qwen3-Coder-Next-bf16) at commit `83d523e8883faaea659705840ce3560472286d08`
using [`mlx-optiq` 0.4.2](https://mlx-optiq.com/docs/cli).
## Attribution and changes
AutomatosX did **not** train Qwen3-Coder-Next. AutomatosX produced this OptiQ
conversion with exact calibration-driven per-layer KL sensitivity:
- candidate bits: 4 and 8
- group size: 64
- target BPW: 5.0 (achieved 5.0235)
- reference: `uniform_4bit` (BF16 reference does not fit in 128 GB UMA for this
80B-class MoE; OptiQ streams BF16 weights for probes against a uniform-4-bit
baseline)
- calibration: 24 脳 512-token sequences from the six-domain OptiQ mix (prose,
thinking, code, agent loops, tool calling, instructions)
- allocation guards: block floor 2, low-bit run limit 3, first/last protection
No new training or benchmark leaderboard claims are made by AutomatosX.
Evaluate coding and tool-use quality on your own tasks.
## Package details
| Property | Value |
| --- | --- |
| Format | MLX Safetensors |
| Architecture | Qwen3NextForCausalLM, mixture of experts |
| Quantization | OptiQ mixed 4/8-bit, group size 64 |
| Measured allocation | 110 components at 4-bit; 431 at 8-bit |
| Target / achieved BPW | 5.0 / 5.0235 |
| Layers / experts | 48 layers 路 512 experts 路 10 active |
| Configured context | 262,144 tokens |
| Logical parameters | 79,674,391,296 (79.67B) |
| Download size | Approximately 55.1 GB |
| Intended hardware | Apple Silicon |
## Download
```bash
hf download AutomatosX/AX-Qwen3-Coder-Next-MLX-OptiQ-4bit \
--local-dir ./AX-Qwen3-Coder-Next-MLX-OptiQ-4bit
```
## Use with MLX-LM
```bash
pip install -U mlx-lm
mlx_lm.generate \
--model AutomatosX/AX-Qwen3-Coder-Next-MLX-OptiQ-4bit \
--prompt "Write a Python function that merges two sorted lists."
```
Apply the included chat template for conversational or tool-using prompts.
## Serve with AX Engine
```bash
ax-engine doctor --mlx-model-artifacts-dir ./AX-Qwen3-Coder-Next-MLX-OptiQ-4bit
ax-engine serve ./AX-Qwen3-Coder-Next-MLX-OptiQ-4bit --port 31418
```
## Validation
Validated on macOS arm64 with mlx-optiq 0.4.2
and AX Engine 6.11.1:
- OptiQ metadata and 4/8-bit allocation: present
- Safetensors index logical parameter override: 79,674,391,296
- MLX-LM load/generation: `passed`
- AX Engine doctor: `ready`
Sample generation (truncated): `"Here's a Python function that merges two sorted lists into a single sorted list:\n\n```python\ndef merge_sorted_lists(list1, list2):\n \"\"\"\n Merge two sorted l"`
## Provenance
Exact conversion inputs, allocation, and weight digests are recorded in
`ax_provenance.json`, `optiq_metadata.json`, and (when present)
`optiq_sensitivity.json`. The source model card is preserved as
`UPSTREAM_README.md`.
## License
Apache License 2.0. See `LICENSE`, the
[Qwen/Qwen3-Coder-Next model card](https://huggingface.co/Qwen/Qwen3-Coder-Next), and
`UPSTREAM_README.md` for limitations and responsible-use guidance.