File size: 1,581 Bytes
c7972e3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
---
license: apache-2.0
language:
- en
tags:
- mlx
- optiq
- diffusion
library_name: mlx
pipeline_tag: text-generation
base_model: inclusionAI/LLaDA2.2-flash
---

# LLaDA2.2-flash-OptiQ-2bit

> **Built with [mlx-optiq](https://mlx-optiq.com)**, the MLX-native toolkit to
> quantize, fine-tune, and serve LLMs locally on Apple Silicon (no PyTorch, no
> cloud). [Try the Lab](https://mlx-optiq.com/docs/lab/) · [All OptiQ
> quants](https://mlx-optiq.com/models) · [Docs](https://mlx-optiq.com/docs/)

An [OptiQ](https://mlx-optiq.com) mixed-precision MLX quant of
**LLaDA2.2-flash**, a ~100B **diffusion** language model with a 256-routed-expert
sparse MoE. This is an extreme 2-bit build.

- **Mixed 2/4-bit `static` build** — per-layer bit-widths assigned to a 2.5
  target bits-per-weight.
- **192 GB bf16 to 36 GB on disk** (5.3x).

## Requirements

Needs `optiq >= 0.4.4`, which ships the vendored `llada2_moe` decoder (the
256-expert diffusion MoE) and the block-diffusion decode loop. Stock `mlx-lm`
has no `llada2_moe` arch and cannot load or generate from this repo.

```bash
pip install -U optiq
```

## Running it

LLaDA2 is a masked-diffusion model, not autoregressive — it denoises a canvas
block by block. `optiq serve` detects the arch and routes it through OptiQ's
vendored decoder and the block-diffusion decode loop automatically:

```bash
optiq serve --model mlx-community/LLaDA2.2-flash-OptiQ-2bit
```

Then call the OpenAI-compatible endpoint at `http://localhost:8000/v1`, or use
it from the [OptiQ Lab](https://mlx-optiq.com/docs/lab/) and `optiq code`.