File size: 6,147 Bytes
de634da
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
dc92f45
de634da
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
dc92f45
de634da
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
Hy3 (hy_v3) for llama.cpp β€” macOS Metal support
Patches and binaries to run Tencent Hy3 (hy_v3) 295B MoE on llama.cpp with Apple Silicon GPU (Metal).

Why this project exists
llama.cpp supports dozens of architectures, but Hy3 (hy_v3) wasn't one of them. Tencent released Hy3 as open-source (Apache 2.0), AngelSlim quantized it to GGUF, but to run it on a Mac someone had to write the missing piece: hy_v3 architecture support in llama.cpp.

This repo contains the patches that add hy_v3 to llama.cpp β€” architecture detection, weight loading, MoE + shared expert forward pass, and MTP self-speculative decoding. All compiled with Metal for Apple Silicon GPU.

The goal is simple: run a 295B model on a MacBook. Not on the cloud, not on a cluster. On a laptop.

Credits
This starts and ends with AngelSlim and their work on HuggingFace:

AngelSlim/Hy3-GGUF β€” GGUF-quantized model (IQ1_M and Q4_K_M), mixed-precision recipes, importance matrix, setup script, benchmarks, chat template. Without this work, this project wouldn't exist.

AngelSlim provided:

The base patches for hy_v3 architecture in llama.cpp
The IQ1_M quantization with mixed recipe (critical weights in Q8_0/Q6_K, experts in IQ1_M/IQ2_XXS)
The importance matrix to allocate bits where they matter
The chat template for tool calling and reasoning
This repo takes those patches, applies them to llama.cpp, and builds them with Metal for macOS.

Thank you AngelSlim. πŸ™Œ

The model
Detail	Value
Architecture	Hy3 (hy_v3) β€” Hunyuan V3
Developed by	Tencent
Parameters	295B
Layers	81 (80 routed + 1 MTP)
Experts	192 (8 active per token)
Gating	Sigmoid + correction bias + top-8 selection
Quantization	IQ1_M (AngelSlim mixed recipe)
File size	~85 GB (with MTP)
Original HF model: Tencent/Hy3 AngelSlim GGUF quant: AngelSlim/Hy3-GGUF Download: Hy3-IQ1_M-mtp.gguf (85 GB, IQ1_M with MTP)

IQ1_M vs BF16 quality loss: ~+0.3% PPL β€” imperceptible. Full benchmarks on AngelSlim's HF page, file assets/benchmark.png.

Why a MacBook?
Mac	RAM	IQ1_M (85 GB)	MTP	Context	Notes
M5 Max	128 GB	βœ…	βœ…	64K	Everything on, comfortable
M4 Max	128 GB	βœ…	βœ…	64K	Everything on
M3 Max	128 GB	βœ…	βœ…	64K	Everything on
MacBook Pro	128 GB	βœ…	βœ…	64K	Runs on a laptop
Mac Studio	96 GB	βœ…	❌	64K	KV q8_0 only, no MTP
MacBook Pro	96 GB	βœ…	❌	64K	Same as above
A 295B MoE running on a MacBook with 128 GB is a concrete milestone for local AI:

No cloud, no API keys, no subscriptions
Total privacy β€” data never leaves your machine
No dedicated GPU needed β€” Apple Silicon unified memory is enough
Portable β€” no server rack, no cluster
With 96 GB it still works: the model is ~85 GB, leaving ~11 GB for the system. Just compress the KV cache (-ctk q8_0 -ctv q8_0) and skip MTP (which adds ~2 GB of weights plus a draft KV cache). With 128 GB everything runs β€” MTP included, with headroom.

Download
1. The GGUF model
# IQ1_M with MTP (85 GB) β€” recommended for 128 GB
wget https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M-mtp.gguf

# IQ1_M without MTP (84 GB) β€” for 96 GB
wget https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M.gguf
2. The code
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
git checkout 19bba67c1
git apply /path/to/0001-add-hyv3-support.patch
cp /path/to/hyv3.cpp src/models/
Or clone the ready-made fork:

git clone https://github.com/<your-username>/llama.cpp-hyv3
cd llama.cpp-hyv3
3. Build
mkdir build && cd build
cmake .. -DLLAMA_METAL=ON
make -j$(sysctl -n hw.logicalcpu)
Commands
CLI (base inference)
./build/bin/llama-cli \
    -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
    -c 65536 \
    -ngl 99 \
    -fa on \
    -ctk q8_0 -ctv q8_0 \
    -p "Hello" \
    -n 100 \
    --temp 0.6
CLI (with reasoning/thinking)
./build/bin/llama-cli \
    -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
    -c 65536 \
    -ngl 99 -fa on \
    -ctk q8_0 -ctv q8_0 \
    -p "Hello" \
    -n 200 \
    --temp 0.6 \
    --reasoning on \
    --reasoning-budget -1

Server (OpenAI-compatible API)
./build/bin/llama-server \
    -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
    -c 65536 \
    -ngl 99 -fa on \
    -ctk q8_0 -ctv q8_0 \
    --temp 0.6 \
    --port 8080
Test API call:

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Hy3-IQ1_M-mtp",
    "messages": [{"role": "user", "content": "Hello"}],
    "temperature": 0.6,
    "max_tokens": 200
  }'

For Mac with 96 GB RAM
# No MTP, compressed KV cache
./build/bin/llama-cli \
    -m ~/Downloads/Hy3-IQ1_M.gguf \
    -c 65536 \
    -ngl 99 -fa on \
    -ctk q8_0 -ctv q8_0 \
    -p "Hello" -n 100 \
    --temp 0.6
Flag reference
Flag	What it does
-m PATH	Path to the GGUF model file
-c N	Context size in tokens. 65536 = 64K. Higher = more memory
-ngl N	Layers to offload to GPU. 99 = all layers on Metal
-fa on	Flash attention β€” reduces memory and speeds up attention
-ctk q8_0 -ctv q8_0	KV cache in q8_0. Essential for 96 GB (saves ~20 GB)
--temp N	Sampling temperature. 0.0 = deterministic/greedy
--reasoning on	Enable thinking/reasoning (tag)
--reasoning-budget N	Max tokens for thinking. -1 = unlimited
--spec-type draft-mtp	MTP self-speculative decoding (*-mtp.gguf only)
--spec-draft-n-max N	Max draft tokens per MTP step
Project structure
β”œβ”€β”€ 0001-add-hyv3-support.patch   # Patch for 9 llama.cpp files (383 lines)
β”œβ”€β”€ src/models/hyv3.cpp           # hy_v3 model implementation + MTP (388 lines)
└── README.md                     # This file
Modified files in llama.cpp
File	Change
src/llama-arch.h	New enum LLM_ARCH_HYV3
src/llama-arch.cpp	Architecture name hy_v3
src/llama-model.cpp	Model mapping + Neox rope type
src/models/models.h	llama_model_hyv3 class declaration
src/models/hyv3.cpp	New β€” load, forward, MTP draft head
gguf-py/gguf/constants.py	Arch enum + tensor list (28 hy_v3 tensors)
gguf-py/gguf/tensor_mapping.py	MTP tensor name mapping
conversion/__init__.py	HF β†’ GGUF model name mapping
common/chat.cpp	Chat template parser (tool calls + reasoning)
License
Apache 2.0. Same as the original Tencent/Hy3 model and AngelSlim's patches.

Long live open local AI. A 295B model running on a MacBook. πŸŽ‰