Instructions to use AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP
Run Hermes
hermes
- OpenClaw new
How to use AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- MLX LM
How to use AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP", "messages": [ {"role": "user", "content": "Hello"} ] }'
AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP
An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved as BF16 sidecars when present.
Development evidence — not a certified AXQuant release. This package has conversion and artifact-integrity records, but it does not publish measured quality, long-context, kernel-speed, or MTP-speed evidence. Do not interpret the AXQ product label as a benchmark claim.
Model details
| Property | Value |
|---|---|
| Base model | Qwen/Qwen3.6-27B |
| Source revision | 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9 |
| Product family | qwen3.6 |
| Source architecture | Qwen3_5ForConditionalGeneration (dense); text path optimized |
| Main-model parameters | 27.36B logical parameters |
| Quantizer | AXQuant 1.0.0 |
| Hub budget class | 4bit |
| AXQuant base precision class | 4bit |
| Planned storage-adjusted BPW | 5.5800 |
| Measured main-model BPW | 5.4183 |
| Measured total BPW, including MTP | 5.5801 |
| Safetensors weight size | 19.38 GB |
| Approximate complete download | 19.40 GB |
| Configured maximum context | 262,144 tokens; practical limits depend on unified memory |
| Primary runtime | AX Engine, compatibility level A |
| Compatible runtime | MLX-LM standard text inference, compatibility level B |
| MTP present | True |
| Vision sidecar present | True |
This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.
Choosing an AXQ pack
AXQ names describe a storage-budget product class, not one uniform precision applied to every
tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative.
In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting
6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection
floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily
protected models.
| Sibling | Intended trade-off |
|---|---|
| 4bit sibling | Lower-storage AXQ budget; check its exact BPW |
| 6bit sibling | Higher average precision near a 6-BPW budget |
See the AutomatosX MLX model catalog for related MLX and OptiQ alternatives.
Download
python -m pip install -U huggingface_hub
hf download AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP --local-dir ./AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP
Allow at least 19.40 GB of free disk space. Pin the resulting Hub commit in reproducible
deployments rather than relying indefinitely on main.
Run with MLX-LM
python -m pip install -U mlx-lm
mlx_lm.generate \
--model AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP \
--prompt "Explain mixed-precision quantization in three sentences." \
--max-tokens 128 \
--temp 0.0
MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime
metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore
does not establish MTP acceleration or vision-language quality. The artifact records MLX
0.32.0 and MLX-LM 0.31.3 from conversion.
Serve with AX Engine and MTP
After installing AX Engine, download the complete repository and serve the local directory:
ax-engine serve ./AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP --port 31418
AX Engine is the authority for the AXQ runtime contract and native MTP sidecar.
This development package does not claim runtime speedups until identical-checkpoint benchmarks are
published. The artifact records AX Engine version not recorded. Native
model-manifest.json status: included as model-manifest.json.
Quantization layout
| Main-weight precision | Parameters | Share |
|---|---|---|
4bit |
24.35B | 87.65% |
8bit |
1.27B | 4.58% |
bf16 |
2.16B | 7.77% |
- Quantization methods:
affine, bf16. - Group sizes used by quantized assignments:
32, 64. - MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
- Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
- Optimization scope:
text-path. - Support tier:
convertible.
BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.
Evidence and validation status
| Check | Status |
|---|---|
| Planning evidence | architecture_prior |
| Calibration | none; the allocation is based on architecture priors |
| Quantizer execution | 497/497 recorded module conversions succeeded; 0 fallbacks |
| AX Engine native manifest | included as model-manifest.json |
| Quality versus BF16 or uniform baselines | Not published; no quality-retention claim |
| MTP acceptance and speed | not measured; no MTP speedup claim |
| AX Engine kernel evidence | unmeasured |
| Vision-language quality | Not evaluated or claimed; vision tensors are preserved at BF16 |
| Long-context quality | 262,144-token capacity is config metadata, not a validated claim |
| Release certification | Not certified; formal AXQuant M0-M8 gates are not closed |
Intended use and limitations
- Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.
- No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.
- Architecture-prior allocation is not measured sensitivity. It must not be presented as measured model quality.
- MTP may be ignored outside AX Engine and its speedup is unmeasured for this exact checkpoint.
- Vision weights are byte-preserved at BF16, but this release does not claim validated VLM quality.
- The configured context window can require substantially more memory as the KV cache grows.
- Upstream capabilities, limitations, biases, and responsible-use guidance still apply.
Provenance and audit files
axquant_manifest.json: package identity, byte accounting, runtime contract, software versions, and file checksums.axquant_plan.json: per-tensor precision decisions and planning evidence.axquant_quantizer_execution.json: conversion coverage and fallback records.axquant_runtime.json: AX Engine and MLX-LM compatibility contract.axquant_mtp_sidecar_manifest.json: MTP tensor provenance.axquant_vision_sidecar_manifest.json: protected vision tensor provenance.model-manifest.json: AX Engine native tensor manifest.
All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. Parallel OptiQ repositories use a different quantizer and should not be assumed to have identical BPW or quality.
License
The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the Qwen/Qwen3.6-27B model card for license terms, model limitations, and responsible-use guidance.
- Downloads last month
- 35
4-bit
Model tree for AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP
Base model
Qwen/Qwen3.6-27B