huckiyang's picture
Update README.md
6b0f47c verified
|
Raw
History Blame Contribute Delete
2.19 kB
---
license: apache-2.0
library_name: mlx
tags:
- mlx
- inkling
- moe
- text-generation
- speech
- vision
- text
- audio
base_model: thinkingmachines/Inkling-NVFP4
pipeline_tag: text-generation
language:
- en
---
# Inkling-mlx (4-bit, text backbone)
An **MLX 4-bit** build of the encoder-free **text backbone** of Thinking Machines' **Inkling**
(975B-total / 41B-active MoE), for running natively on Apple Silicon with
[`mlx-lm`](https://github.com/ml-explore/mlx-lm).
This is created for people using a two Apple Mac Studio M3 Ultra with 192/512 GB. (for one Mac Studio you need 2-bit!)
> **Community note** This is **not fullly numerically-verified**
> conversion, shared to see whether anyone can load/run a model this large on Apple
> Silicon and to gather feedback. Expect rough edges; please open a discussion with
> results (or failures).
## Notes
- **Memory:** the build is ~**580 GB** on disk (4-bit routed experts + bf16 attention /
shared experts / embeddings). Loading it needs roughly that much **unified memory**, beyond
any single Mac today (max 512 GB), so realistically it needs distributed/multi-device MLX or
a very large box. This is largely a **research artifact**.
- **Not fully verified yet:** the custom Inkling forward (factorized attention + short-conv + sigmoid
MoE) is a from-reference reimplementation whose logits have **not** been checked against the
original. Correctness is unconfirmed.
## Provenance
- **Source:** `thinkingmachines/Inkling-NVFP4` (NVFP4) → dequantized → MLX affine **4-bit**
(group size 64). Only the routed MoE experts are quantized; everything else is bf16.
- NVFP4→INT4 re-quantization, so expect small extra quality loss vs. a BF16-sourced build.
## Usage (once a loader is available)
```python
from mlx_lm import load, generate
model, tokenizer = load("mlx-community/Inkling-NVFP4-mlx-4bit")
print(generate(model, tokenizer, prompt="The capital of France is", max_tokens=64))
```
> The custom model class lives in the conversion repo (`models/inkling_mlx.py`); until it's
> registered in `mlx-lm`, load via that module's `load()`.
>
> Blog: https://huckiyang.github.io/blog/inkling-audio-design.html