AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP

An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved as BF16 sidecars when present.

Development evidence — not a certified AXQuant release. This package has conversion and artifact-integrity records, but it does not publish measured quality, long-context, kernel-speed, or MTP-speed evidence. Do not interpret the AXQ product label as a benchmark claim.

Model details

Property Value
Base model Qwen/Qwen3.6-27B
Source revision 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Product family qwen3.6
Source architecture Qwen3_5ForConditionalGeneration (dense); text path optimized
Main-model parameters 27.36B logical parameters
Quantizer AXQuant 1.0.0
Hub budget class 4bit
AXQuant base precision class 4bit
Planned storage-adjusted BPW 5.5800
Measured main-model BPW 5.4183
Measured total BPW, including MTP 5.5801
Safetensors weight size 19.38 GB
Approximate complete download 19.40 GB
Configured maximum context 262,144 tokens; practical limits depend on unified memory
Primary runtime AX Engine, compatibility level A
Compatible runtime MLX-LM standard text inference, compatibility level B
MTP present True
Vision sidecar present True

This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.

Choosing an AXQ pack

AXQ names describe a storage-budget product class, not one uniform precision applied to every tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily protected models.

Sibling Intended trade-off
4bit sibling Lower-storage AXQ budget; check its exact BPW
6bit sibling Higher average precision near a 6-BPW budget

See the AutomatosX MLX model catalog for related MLX and OptiQ alternatives.

Download

python -m pip install -U huggingface_hub
hf download AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP --local-dir ./AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP

Allow at least 19.40 GB of free disk space. Pin the resulting Hub commit in reproducible deployments rather than relying indefinitely on main.

Run with MLX-LM

python -m pip install -U mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP \
  --prompt "Explain mixed-precision quantization in three sentences." \
  --max-tokens 128 \
  --temp 0.0

MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore does not establish MTP acceleration or vision-language quality. The artifact records MLX 0.32.0 and MLX-LM 0.31.3 from conversion.

Serve with AX Engine and MTP

After installing AX Engine, download the complete repository and serve the local directory:

ax-engine serve ./AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP --port 31418

AX Engine is the authority for the AXQ runtime contract and native MTP sidecar. This development package does not claim runtime speedups until identical-checkpoint benchmarks are published. The artifact records AX Engine version not recorded. Native model-manifest.json status: included as model-manifest.json.

Quantization layout

Main-weight precision Parameters Share
4bit 24.35B 87.65%
8bit 1.27B 4.58%
bf16 2.16B 7.77%
  • Quantization methods: affine, bf16.
  • Group sizes used by quantized assignments: 32, 64.
  • MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
  • Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
  • Optimization scope: text-path.
  • Support tier: convertible.

BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.

Evidence and validation status

Check Status
Planning evidence architecture_prior
Calibration none; the allocation is based on architecture priors
Quantizer execution 497/497 recorded module conversions succeeded; 0 fallbacks
AX Engine native manifest included as model-manifest.json
Quality versus BF16 or uniform baselines Not published; no quality-retention claim
MTP acceptance and speed not measured; no MTP speedup claim
AX Engine kernel evidence unmeasured
Vision-language quality Not evaluated or claimed; vision tensors are preserved at BF16
Long-context quality 262,144-token capacity is config metadata, not a validated claim
Release certification Not certified; formal AXQuant M0-M8 gates are not closed

Intended use and limitations

  • Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.
  • No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.
  • Architecture-prior allocation is not measured sensitivity. It must not be presented as measured model quality.
  • MTP may be ignored outside AX Engine and its speedup is unmeasured for this exact checkpoint.
  • Vision weights are byte-preserved at BF16, but this release does not claim validated VLM quality.
  • The configured context window can require substantially more memory as the KV cache grows.
  • Upstream capabilities, limitations, biases, and responsible-use guidance still apply.

Provenance and audit files

All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. Parallel OptiQ repositories use a different quantizer and should not be assumed to have identical BPW or quality.

License

The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the Qwen/Qwen3.6-27B model card for license terms, model limitations, and responsible-use guidance.

Downloads last month
35
Safetensors
Model size
27B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP

Base model

Qwen/Qwen3.6-27B
Quantized
(681)
this model

Collection including AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP