File size: 4,880 Bytes
2baac9f
 
c4c9f21
 
8a0debe
c4c9f21
2baac9f
 
 
8a0debe
2baac9f
c4c9f21
 
 
 
 
 
 
2baac9f
 
8a0debe
2baac9f
8a0debe
 
 
 
 
2baac9f
8a0debe
 
c4c9f21
8a0debe
 
 
 
 
43c9020
c4c9f21
2baac9f
 
 
c4c9f21
 
 
 
2baac9f
c4c9f21
 
2baac9f
c4c9f21
 
2baac9f
c4c9f21
2baac9f
c4c9f21
 
 
 
 
 
 
 
 
 
 
2baac9f
c4c9f21
2baac9f
c4c9f21
2baac9f
 
 
c4c9f21
2baac9f
af0a81b
c4c9f21
 
 
af0a81b
2baac9f
 
c4c9f21
 
 
2baac9f
c4c9f21
2baac9f
 
 
 
af0a81b
c4c9f21
 
2baac9f
 
 
c4c9f21
 
2baac9f
 
 
 
 
c4c9f21
 
 
 
2baac9f
 
 
c4c9f21
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
---
license: apache-2.0
license_link: https://www.apache.org/licenses/LICENSE-2.0
thumbnail: https://huggingface.co/AtomicChat/Inkling-GGUF/resolve/main/hero.png
base_model:
- thinkingmachines/Inkling
base_model_relation: quantized
quantized_by: AtomicChat
pipeline_tag: text-generation
library_name: gguf
tags:
- atomic-chat
- inkling
- thinkingmachines
- gguf
- llama.cpp
- imatrix
- quantized
---

<center>

<div style="display:flex; justify-content:center; align-items:center; gap:2%; max-width:560px; margin:0 auto;">
<a href="https://atomic.chat" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/Inkling-GGUF/resolve/main/pill_atomic_v3.png" alt="Atomic Chat" style="width:100%; height:auto; max-width:186px;"></a>
<a href="https://discord.gg/8wGSsvmg4V" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/Inkling-GGUF/resolve/main/pill_discord_v3.png" alt="Join Discord" style="width:100%; height:auto; max-width:184px;"></a>
<a href="https://github.com/AtomicBot-ai/Atomic-Chat" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/Inkling-GGUF/resolve/main/pill_github_v3.png" alt="GitHub" style="width:100%; height:auto; max-width:141px;"></a>
</div>

<br/>

<img src="https://huggingface.co/AtomicChat/Inkling-GGUF/resolve/main/hero.png" alt="Inkling" style="width:100%; max-width:100%; height:auto; margin-bottom:0.6em;"/>

<div style="display:flex; justify-content:center; gap:0.5em;">
<a href="https://huggingface.co/thinkingmachines/Inkling"><strong>Base model: thinkingmachines/Inkling</strong></a>
</div>
</center>

**Inkling**, self-quantized to GGUF by [Atomic Chat](https://atomic.chat). Built straight from Thinking Machines Lab's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.

## Highlights

- **952.4B parameters**: the weights this repo quantizes.
- **66 layers**: Mixture-of-Experts.
- **Modalities**: the base model handles Text, Image, Audio; this repo ships text-only quants, it carries no vision projector.
- **Full imatrix ladder**: every quant is calibrated with an importance matrix, published here alongside the quants.

> [!NOTE]
> These GGUFs are **self-quantized from the original weights**, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.

> [!IMPORTANT]
> Always pass `--jinja` so the **Inkling chat template** is applied. Without it the model can emit malformed turns.

## Model Overview

| Property | Value |
|---|---|
| Base model | `thinkingmachines/Inkling` |
| Parameters | 952.4B |
| Layers | 66 |
| Experts | 256 routed (top-6) |
| Context length | not stated |
| Vocabulary | 201,024 |
| Modalities | Text, Image, Audio in the base model; text only in this repo, it ships no vision projector |
| Architecture | Mixture-of-Experts, 256 experts (top-6), 64 attention heads over 8 KV heads, `InklingForConditionalGeneration` |
| This repo | GGUF quants (imatrix); the importance matrix is published here as `imatrix/imatrix-code-at_128.gguf` |

<img src="https://huggingface.co/AtomicChat/Inkling-GGUF/resolve/main/benchmark.png" alt="Inkling benchmark scores" style="width:100%; max-width:900px;"/>

Scores are Thinking Machines Lab's published results for the base `thinkingmachines/Inkling`, not our own measurements. Quantization preserves the large majority of this; `Q4_K_M` and up stay close to full precision.

## Get started

Run Inkling locally with:

- **[Atomic Chat](https://atomic.chat):** the easiest path. Open the app, search `AtomicChat/Inkling-GGUF`, pick a quant, hit **Use this model**.
- **llama.cpp:** `llama-server -hf AtomicChat/Inkling-GGUF:None --jinja -c 8192`
- **Ollama:** `ollama run hf.co/AtomicChat/Inkling-GGUF:None`
- **LM Studio / Jan:** search the repo id, download any quant.

## Best practices

| Parameter | Value |
|---|---|
| sampling defaults | not stated |

The base model card does not state sampling defaults.

## Run in llama.cpp

```bash
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --target llama-cli llama-server
```

```bash
./llama.cpp/build/bin/llama-server \
    -hf AtomicChat/Inkling-GGUF:None \
    --jinja -ngl 99 -c 8192 -fa on
```

## How these were made

1. Download `thinkingmachines/Inkling` (original weights).
2. Convert to f16 GGUF with [llama.cpp](https://github.com/ggml-org/llama.cpp).
3. Build an importance matrix over our calibration corpus, published here as `imatrix/imatrix-code-at_128.gguf`.
4. Quantize the ladder with `--imatrix`.

## License

Original model by Thinking Machines Lab, released under the Apache 2.0 license. Full terms: [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0). Quantized by Atomic Chat.