Text Generation
MLX
Safetensors
qwen3_next
apple-silicon
quantized
mixed-precision
axquant
axq
development
qwen3-next
4bit
4-bit precision
conversational
Instructions to use AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit
Run Hermes
hermes
- OpenClaw new
How to use AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- MLX LM
How to use AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }'
| license: apache-2.0 | |
| library_name: mlx | |
| base_model: Qwen/Qwen3-Coder-Next | |
| base_model_relation: quantized | |
| pipeline_tag: text-generation | |
| tags: | |
| - mlx | |
| - apple-silicon | |
| - quantized | |
| - mixed-precision | |
| - axquant | |
| - axq | |
| - development | |
| - qwen3-next | |
| - 4bit | |
| - 4-bit | |
| # AX-Qwen3-Coder-Next-MLX-AXQ-4bit | |
| An **AXQuant (AXQ)** mixed-precision MLX checkpoint for Apple Silicon, converted directly from | |
| the BF16 source model. The language path is quantized under AXQuant protection floors (embeddings, norms, and other protected tensors remain higher precision). | |
| > **Development evidence — not a certified AXQuant release.** This package has conversion and | |
| > artifact-integrity records, but it does not publish measured quality, long-context, kernel-speed, | |
| > or MTP-speed evidence. Do not interpret the AXQ product label as a benchmark claim. | |
| ## Model details | |
| | Property | Value | | |
| | --- | --- | | |
| | Base model | [Qwen/Qwen3-Coder-Next](https://huggingface.co/Qwen/Qwen3-Coder-Next) | | |
| | Source revision | `unrecorded` | | |
| | Product family | `qwen3-next` | | |
| | Source architecture | `Qwen3NextForCausalLM` (mixture of experts (MoE)); text path optimized | | |
| | Main-model parameters | 79.67B logical parameters | | |
| | Quantizer | AXQuant `1.0.1` | | |
| | Hub budget class | `4bit` | | |
| | AXQuant base precision class | `4bit` | | |
| | Planned storage-adjusted BPW | 15.7300 | | |
| | Measured main-model BPW | 15.7300 | | |
| | Measured total BPW | **15.7300** | | |
| | Safetensors weight size | 156.66 GB | | |
| | Approximate complete download | 156.67 GB | | |
| | Configured maximum context | 262,144 tokens; practical limits depend on unified memory | | |
| | Primary runtime | AX Engine, compatibility level A | | |
| | Compatible runtime | MLX-LM standard text inference, compatibility level B | | |
| | MTP present | `False` | | |
| | Vision sidecar present | `False` | | |
| This repository contains MLX Safetensors. It does **not** contain PyTorch or GGUF weights. | |
| ## Choosing an AXQ pack | |
| AXQ names describe a **storage-budget product class**, not one uniform precision applied to every | |
| tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. | |
| In particular, a `6bit`-named mixed plan may retain `4bit` as its base precision while selecting | |
| 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection | |
| floors can also raise a `4bit`-named pack close to (or above) a `6bit` budget on small or heavily | |
| protected models. | |
| | Sibling | Intended trade-off | | |
| | --- | --- | | |
| | [4bit sibling](https://huggingface.co/AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit) | Lower-storage AXQ budget; check its exact BPW | | |
| | [6bit sibling](https://huggingface.co/AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-6bit) | Higher average precision near a 6-BPW budget | | |
| See the [AutomatosX MLX model catalog](https://huggingface.co/collections/AutomatosX/automatosx-mlx-model-catalog) | |
| for related MLX and OptiQ alternatives. | |
| ## Download | |
| ```bash | |
| python -m pip install -U huggingface_hub | |
| hf download AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit --local-dir ./AX-Qwen3-Coder-Next-MLX-AXQ-4bit | |
| ``` | |
| Allow at least 156.67 GB of free disk space. Pin the resulting Hub commit in reproducible | |
| deployments rather than relying indefinitely on `main`. | |
| ## Run with MLX-LM | |
| ```bash | |
| python -m pip install -U mlx-lm | |
| mlx_lm.generate \ | |
| --model AutomatosX/AX-Qwen3-Coder-Next-MLX-AXQ-4bit \ | |
| --prompt "Explain mixed-precision quantization in three sentences." \ | |
| --max-tokens 128 \ | |
| --temp 0.0 | |
| ``` | |
| MLX-LM compatibility covers standard **text/backbone inference**. It may ignore AXQuant runtime | |
| metadata and optional sidecars (`vision.safetensors`, `mtp.safetensors`); this command therefore | |
| does not establish MTP acceleration or vision-language quality. The artifact records MLX | |
| `0.32.0` and MLX-LM `0.31.3` from conversion. | |
| ## Serve with AX Engine | |
| After installing [AX Engine](https://github.com/defai-digital/ax-engine), download the complete | |
| repository and serve the local directory: | |
| ```bash | |
| ax-engine serve ./AX-Qwen3-Coder-Next-MLX-AXQ-4bit --port 31418 | |
| ``` | |
| AX Engine is the authority for the AXQ runtime contract. | |
| This development package does not claim runtime speedups until identical-checkpoint benchmarks are | |
| published. The artifact records AX Engine version `not recorded`. Native | |
| `model-manifest.json` status: not included. | |
| ## Quantization layout | |
| | Main-weight precision | Parameters | Share | | |
| | --- | ---: | ---: | | |
| | `4bit` | 1.69B | 2.12% | | |
| | `6bit` | 397,312 | 0.00% | | |
| | `8bit` | 361.50M | 0.45% | | |
| | `bf16` | 77.62B | 97.42% | | |
| - Quantization methods: `affine, bf16`. | |
| - Group sizes used by quantized assignments: `32, 64`. | |
| - MTP sidecar: not included. | |
| - Vision sidecar: not included. | |
| - Optimization scope: `text-path`. | |
| - Support tier: `convertible`. | |
| BF16 sidecars, when present, are included in total download size. Their presence does not by itself | |
| establish MTP acceleration or vision-language quality. | |
| ## Evidence and validation status | |
| | Check | Status | | |
| | --- | --- | | |
| | Planning evidence | `architecture_prior` | | |
| | Calibration | none; the allocation is based on architecture priors | | |
| | Quantizer execution | 397/397 recorded module conversions succeeded; 0 fallbacks | | |
| | AX Engine native manifest | not included | | |
| | Quality versus BF16 or uniform baselines | Not published; no quality-retention claim | | |
| | MTP acceptance and speed | not measured; no MTP speedup claim | | |
| | AX Engine kernel evidence | `unmeasured` | | |
| | Vision-language quality | Not applicable (no vision sidecar in this package) | | |
| | Long-context quality | 262,144-token capacity is config metadata, not a validated claim | | |
| | Release certification | **Not certified**; formal AXQuant M0-M8 gates are not closed | | |
| ## Intended use and limitations | |
| - Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes. | |
| - No minimum unified-memory figure is claimed; loadability depends on model size, context length, | |
| KV-cache policy, runtime buffers, and other processes using unified memory. | |
| - Architecture-prior allocation is not measured sensitivity. It must not be presented as measured | |
| model quality. | |
| - The configured context window can require substantially more memory as the KV cache grows. | |
| - Upstream capabilities, limitations, biases, and responsible-use guidance still apply. | |
| ## Provenance and audit files | |
| - [`axquant_manifest.json`](axquant_manifest.json): package identity, byte accounting, runtime | |
| contract, software versions, and file checksums. | |
| - [`axquant_plan.json`](axquant_plan.json): per-tensor precision decisions and planning evidence. | |
| - [`axquant_quantizer_execution.json`](axquant_quantizer_execution.json): conversion coverage and | |
| fallback records. | |
| - [`axquant_runtime.json`](axquant_runtime.json): AX Engine and MLX-LM compatibility contract. | |
| All published provenance uses repository-relative paths. Local source paths are stripped before | |
| publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ | |
| artifact. Parallel OptiQ repositories use a different quantizer and should not be assumed to have | |
| identical BPW or quality. | |
| ## License | |
| The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See | |
| the [Qwen/Qwen3-Coder-Next model card](https://huggingface.co/Qwen/Qwen3-Coder-Next) for license terms, model | |
| limitations, and responsible-use guidance. | |