File size: 8,077 Bytes
d3ff80a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
# Hy3 (hy_v3) for llama.cpp β€” macOS Metal support

Patches and binaries to run **Tencent Hy3 (hy_v3) 295B MoE** on **llama.cpp** with **Apple Silicon GPU (Metal)**.

---

## Why this project exists

llama.cpp supports dozens of architectures, but **Hy3 (hy_v3) wasn't one of them**. Tencent released Hy3 as open-source (Apache 2.0), AngelSlim quantized it to GGUF, but to run it on a Mac someone had to write the missing piece: hy_v3 architecture support in llama.cpp.

This repo contains **the patches that add hy_v3 to llama.cpp** β€” architecture detection, weight loading, MoE + shared expert forward pass, and MTP self-speculative decoding. All compiled with **Metal** for Apple Silicon GPU.

The goal is simple: **run a 295B model on a MacBook**. Not on the cloud, not on a cluster. On a laptop.

---

## Credits

This starts and ends with **AngelSlim** and their work on HuggingFace:

[**AngelSlim/Hy3-GGUF**](https://huggingface.co/AngelSlim/Hy3-GGUF) β€” GGUF-quantized model (IQ1_M and Q4_K_M), mixed-precision recipes, importance matrix, setup script, benchmarks, chat template. Without this work, this project wouldn't exist.

AngelSlim provided:
- The **base patches** for hy_v3 architecture in llama.cpp
- The **IQ1_M quantization** with mixed recipe (critical weights in Q8_0/Q6_K, experts in IQ1_M/IQ2_XXS)
- The **importance matrix** to allocate bits where they matter
- The **chat template** for tool calling and reasoning

This repo takes those patches, applies them to llama.cpp, and **builds them with Metal for macOS**.

**Thank you AngelSlim.** πŸ™Œ

---

## The model

| Detail | Value |
|--------|-------|
| **Architecture** | Hy3 (hy_v3) β€” Hunyuan V3 |
| **Developed by** | Tencent |
| **Parameters** | 295B |
| **Layers** | 81 (80 routed + 1 MTP) |
| **Experts** | 192 (8 active per token) |
| **Gating** | Sigmoid + correction bias + top-8 selection |
| **Quantization** | IQ1_M (AngelSlim mixed recipe) |
| **File size** | ~85 GB (with MTP) |

> Original HF model: [Tencent/Hy3](https://huggingface.co/tencent/Hy3)
> AngelSlim GGUF quant: [**AngelSlim/Hy3-GGUF**](https://huggingface.co/AngelSlim/Hy3-GGUF)
> Download: [Hy3-IQ1_M-mtp.gguf](https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M-mtp.gguf) (85 GB, IQ1_M with MTP)

**IQ1_M vs BF16 quality loss**: ~+0.3% PPL β€” **imperceptible**. Full benchmarks on [AngelSlim's HF page](https://huggingface.co/AngelSlim/Hy3-GGUF), file `assets/benchmark.png`.

---

## Why a MacBook?

| Mac | RAM | IQ1_M (85 GB) | MTP | Context | Notes |
|-----|-----|:---:|:---:|:--------:|-------|
| **M5 Max** | 128 GB | βœ… | βœ… | 64K | Everything on, comfortable |
| **M4 Max** | 128 GB | βœ… | βœ… | 64K | Everything on |
| **M3 Max** | 128 GB | βœ… | βœ… | 64K | Everything on |
| **MacBook Pro** | 128 GB | βœ… | βœ… | 64K | Runs on a laptop |
| **Mac Studio** | 96 GB | βœ… | ❌ | 64K | KV q8_0 only, no MTP |
| **MacBook Pro** | 96 GB | βœ… | ❌ | 64K | Same as above |

A **295B MoE running on a MacBook with 128 GB** is a concrete milestone for local AI:

- **No cloud**, no API keys, no subscriptions
- **Total privacy** β€” data never leaves your machine
- **No dedicated GPU needed** β€” Apple Silicon unified memory is enough
- **Portable** β€” no server rack, no cluster

With 96 GB it still works: the model is ~85 GB, leaving ~11 GB for the system. Just compress the KV cache (`-ctk q8_0 -ctv q8_0`) and skip MTP (which adds ~2 GB of weights plus a draft KV cache). With 128 GB everything runs β€” MTP included, with headroom.

---

## Download

### 1. The GGUF model

```bash
# IQ1_M with MTP (85 GB) β€” recommended for 128 GB
wget https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M-mtp.gguf

# IQ1_M without MTP (84 GB) β€” for 96 GB
wget https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M.gguf
```

### 2. The code

```bash
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
git checkout 19bba67c1
git apply /path/to/0001-add-hyv3-support.patch
cp /path/to/hyv3.cpp src/models/
```

Or clone the ready-made fork:

```bash
git clone https://github.com/RobZombAI/llama.cpp-metal_hyv3
cd llama.cpp-metal_hyv3
```

### 3. Build

```bash
mkdir build && cd build
cmake .. -DLLAMA_METAL=ON
make -j$(sysctl -n hw.logicalcpu)
```

---

## Commands

### CLI (base inference)

```bash
./build/bin/llama-cli \
    -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
    -c 65536 \
    -ngl 99 \
    -fa on \
    -ctk q8_0 -ctv q8_0 \
    -p "Hello" \
    -n 100 \
    --temp 0.6
```

### CLI (with reasoning/thinking)

```bash
./build/bin/llama-cli \
    -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
    -c 65536 \
    -ngl 99 -fa on \
    -ctk q8_0 -ctv q8_0 \
    -p "Hello" \
    -n 200 \
    --temp 0.6 \
    --reasoning on \
    --reasoning-budget -1
```

### CLI (with MTP self-speculative β€” higher throughput)

```bash
./build/bin/llama-cli \
    -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
    -c 65536 \
    -ngl 99 -fa on \
    --spec-type draft-mtp \
    --spec-draft-n-max 3 \
    --spec-draft-n-min 1 \
    -ctk q8_0 -ctv q8_0 \
    -ctkd q8_0 -ctvd q8_0 \
    -p "Hello" \
    -n 200 \
    --temp 0.6
```

### Server (OpenAI-compatible API)

```bash
./build/bin/llama-server \
    -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
    -c 65536 \
    -ngl 99 -fa on \
    -ctk q8_0 -ctv q8_0 \
    --temp 0.6 \
    --port 8080
```

Test API call:

```bash
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Hy3-IQ1_M-mtp",
    "messages": [{"role": "user", "content": "Hello"}],
    "temperature": 0.6,
    "max_tokens": 200
  }'
```

### Server (with reasoning + MTP)

```bash
./build/bin/llama-server \
    -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
    -c 65536 \
    -ngl 99 -fa on \
    --spec-type draft-mtp \
    --spec-draft-n-max 3 \
    --spec-draft-n-min 1 \
    -ctk q8_0 -ctv q8_0 \
    -ctkd q8_0 -ctvd q8_0 \
    --temp 0.6 \
    --reasoning on \
    --reasoning-budget -1 \
    --port 8080
```

### For Mac with 96 GB RAM

```bash
# No MTP, compressed KV cache
./build/bin/llama-cli \
    -m ~/Downloads/Hy3-IQ1_M.gguf \
    -c 65536 \
    -ngl 99 -fa on \
    -ctk q8_0 -ctv q8_0 \
    -p "Hello" -n 100 \
    --temp 0.6
```

---

## Flag reference

| Flag | What it does |
|------|--------------|
| `-m PATH` | Path to the GGUF model file |
| `-c N` | Context size in tokens. `65536` = 64K. Higher = more memory |
| `-ngl N` | Layers to offload to GPU. `99` = all layers on Metal |
| `-fa on` | Flash attention β€” reduces memory and speeds up attention |
| `-ctk q8_0 -ctv q8_0` | KV cache in q8_0. **Essential for 96 GB** (saves ~20 GB) |
| `--temp N` | Sampling temperature. `0.0` = deterministic/greedy |
| `--reasoning on` | Enable thinking/reasoning (tag) |
| `--reasoning-budget N` | Max tokens for thinking. `-1` = unlimited |
| `--spec-type draft-mtp` | MTP self-speculative decoding (*-mtp.gguf only) |
| `--spec-draft-n-max N` | Max draft tokens per MTP step |

---

## Project structure

```
β”œβ”€β”€ 0001-add-hyv3-support.patch   # Patch for 9 llama.cpp files (383 lines)
β”œβ”€β”€ src/models/hyv3.cpp           # hy_v3 model implementation + MTP (388 lines)
└── README.md                     # This file
```

## Modified files in llama.cpp

| File | Change |
|------|--------|
| `src/llama-arch.h` | New enum `LLM_ARCH_HYV3` |
| `src/llama-arch.cpp` | Architecture name `hy_v3` |
| `src/llama-model.cpp` | Model mapping + Neox rope type |
| `src/models/models.h` | `llama_model_hyv3` class declaration |
| `src/models/hyv3.cpp` | **New** β€” load, forward, MTP draft head |
| `gguf-py/gguf/constants.py` | Arch enum + tensor list (28 hy_v3 tensors) |
| `gguf-py/gguf/tensor_mapping.py` | MTP tensor name mapping |
| `conversion/__init__.py` | HF β†’ GGUF model name mapping |
| `common/chat.cpp` | Chat template parser (tool calls + reasoning) |

## License

**Apache 2.0.** Same as the original [Tencent/Hy3](https://huggingface.co/tencent/Hy3) model and [AngelSlim](https://huggingface.co/AngelSlim)'s patches.

---

**Long live open local AI. A 295B model running on a MacBook. πŸŽ‰**