AX-DeepSeek-V4-Flash-0731-MLX-AXQ-2bit

Development / experimental AXQuant 2-bit pack of deepseek-ai/DeepSeek-V4-Flash-0731 @ 7872f01b1d1fe23eabc4c98b48bffcef5a386062.

Converted on df-macstudio-m2 (Apple M2 Ultra, 192 GB) from the native FP8 0731 source (quant_method=fp8). Product class 2bit-experimental.

This is not the older DeepSeek-V4-Flash Hub pack. Do not treat AX-DeepSeek-V4-Flash-MLX-AXQ-2bit certificates as evidence for this 0731 revision.

Measured precision

Property Value
Target class 2bit-experimental
Measured main BPW 3.1328993873020314
Measured total BPW 3.2142055528774454
Weight bytes 122,212,298,775
Source deepseek-ai/DeepSeek-V4-Flash-0731@7872f01b1d1fe23eabc4c98b48bffcef5a386062
Convert host df-macstudio-m2
AXQuant 1.8.1
mlx / mlx-lm 0.32.0 / 0.31.3 (vendored deepseek_v4 + FP8 load hook)

Claims

Claim Status
Converted on Studio from the pinned 0731 revision Yes
mlx-lm load + generate smoke Passed on df-macstudio-m2
Official DSV4 chat_template.jinja In pack
Checkpoint Tier 1 (generation viability suite) Not certified on this record; Studio 15+15 factory QA with chat is 0.633 combined
AX Engine native manifest Not generatedgenerate-manifest --validate rejected split switch_mlp.gate_proj / up_proj ([256, 2048, 256] vs fused [256, 4096, 256])
MTP acceleration Not certified

Requires AX_ENGINE_2BIT_EXPERIMENTAL=1 if served with AX Engine after a future manifest fix. mlx-lm generate does not need that env.

Attribution

Base weights © DeepSeek. Quantization by AXQuant (development).

Downloads last month
132
Safetensors
Model size
38B params
Tensor type
BF16
·
F32
·
U32
·
I32
·
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AutomatosX/AX-DeepSeek-V4-Flash-0731-MLX-AXQ-2bit

Quantized
(155)
this model

Collections including AutomatosX/AX-DeepSeek-V4-Flash-0731-MLX-AXQ-2bit