Hy3 (hy_v3) for llama.cpp β macOS Metal support
Patches and binaries to run Tencent Hy3 (hy_v3) 295B MoE on llama.cpp with Apple Silicon GPU (Metal).
Why this project exists
llama.cpp supports dozens of architectures, but Hy3 (hy_v3) wasn't one of them. Tencent released Hy3 as open-source (Apache 2.0), AngelSlim quantized it to GGUF, but to run it on a Mac someone had to write the missing piece: hy_v3 architecture support in llama.cpp.
This repo contains the patches that add hy_v3 to llama.cpp β architecture detection, weight loading, MoE + shared expert forward pass, and MTP self-speculative decoding. All compiled with Metal for Apple Silicon GPU.
The goal is simple: run a 295B model on a MacBook. Not on the cloud, not on a cluster. On a laptop.
Credits
This starts and ends with AngelSlim and their work on HuggingFace:
AngelSlim/Hy3-GGUF β GGUF-quantized model (IQ1_M and Q4_K_M), mixed-precision recipes, importance matrix, setup script, benchmarks, chat template. Without this work, this project wouldn't exist.
AngelSlim provided:
- The base patches for hy_v3 architecture in llama.cpp
- The IQ1_M quantization with mixed recipe (critical weights in Q8_0/Q6_K, experts in IQ1_M/IQ2_XXS)
- The importance matrix to allocate bits where they matter
- The chat template for tool calling and reasoning
This repo takes those patches, applies them to llama.cpp, and builds them with Metal for macOS.
Thank you AngelSlim. π
The model
| Detail | Value |
|---|---|
| Architecture | Hy3 (hy_v3) β Hunyuan V3 |
| Developed by | Tencent |
| Parameters | 295B |
| Layers | 81 (80 routed + 1 MTP) |
| Experts | 192 (8 active per token) |
| Gating | Sigmoid + correction bias + top-8 selection |
| Quantization | IQ1_M (AngelSlim mixed recipe) |
| File size | ~85 GB (with MTP) |
Original HF model: Tencent/Hy3 AngelSlim GGUF quant: AngelSlim/Hy3-GGUF Download: Hy3-IQ1_M-mtp.gguf (85 GB, IQ1_M with MTP)
IQ1_M vs BF16 quality loss: ~+0.3% PPL β imperceptible. Full benchmarks on AngelSlim's HF page, file assets/benchmark.png.
Why a MacBook?
| Mac | RAM | IQ1_M (85 GB) | MTP | Context | Notes |
|---|---|---|---|---|---|
| M5 Max | 128 GB | β | β | 64K | Everything on, comfortable |
| M4 Max | 128 GB | β | β | 64K | Everything on |
| M3 Max | 128 GB | β | β | 64K | Everything on |
| MacBook Pro | 128 GB | β | β | 64K | Runs on a laptop |
| Mac Studio | 96 GB | β | β | 64K | KV q8_0 only, no MTP |
| MacBook Pro | 96 GB | β | β | 64K | Same as above |
A 295B MoE running on a MacBook with 128 GB is a concrete milestone for local AI:
- No cloud, no API keys, no subscriptions
- Total privacy β data never leaves your machine
- No dedicated GPU needed β Apple Silicon unified memory is enough
- Portable β no server rack, no cluster
With 96 GB it still works: the model is ~85 GB, leaving ~11 GB for the system. Just compress the KV cache (-ctk q8_0 -ctv q8_0) and skip MTP (which adds ~2 GB of weights plus a draft KV cache). With 128 GB everything runs β MTP included, with headroom.
Download
1. The GGUF model
# IQ1_M with MTP (85 GB) β recommended for 128 GB
wget https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M-mtp.gguf
# IQ1_M without MTP (84 GB) β for 96 GB
wget https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M.gguf
2. The code
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
git checkout 19bba67c1
git apply /path/to/0001-add-hyv3-support.patch
cp /path/to/hyv3.cpp src/models/
Or clone the ready-made fork:
git clone https://github.com/RobZombAI/llama.cpp-metal_hyv3
cd llama.cpp-metal_hyv3
3. Build
mkdir build && cd build
cmake .. -DLLAMA_METAL=ON
make -j$(sysctl -n hw.logicalcpu)
Commands
CLI (base inference)
./build/bin/llama-cli \
-m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
-c 65536 \
-ngl 99 \
-fa on \
-ctk q8_0 -ctv q8_0 \
-p "Hello" \
-n 100 \
--temp 0.6
CLI (with reasoning/thinking)
./build/bin/llama-cli \
-m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
-c 65536 \
-ngl 99 -fa on \
-ctk q8_0 -ctv q8_0 \
-p "Hello" \
-n 200 \
--temp 0.6 \
--reasoning on \
--reasoning-budget -1
CLI (with MTP self-speculative β higher throughput)
./build/bin/llama-cli \
-m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
-c 65536 \
-ngl 99 -fa on \
--spec-type draft-mtp \
--spec-draft-n-max 3 \
--spec-draft-n-min 1 \
-ctk q8_0 -ctv q8_0 \
-ctkd q8_0 -ctvd q8_0 \
-p "Hello" \
-n 200 \
--temp 0.6
Server (OpenAI-compatible API)
./build/bin/llama-server \
-m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
-c 65536 \
-ngl 99 -fa on \
-ctk q8_0 -ctv q8_0 \
--temp 0.6 \
--port 8080
Test API call:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Hy3-IQ1_M-mtp",
"messages": [{"role": "user", "content": "Hello"}],
"temperature": 0.6,
"max_tokens": 200
}'
Server (with reasoning + MTP)
./build/bin/llama-server \
-m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
-c 65536 \
-ngl 99 -fa on \
--spec-type draft-mtp \
--spec-draft-n-max 3 \
--spec-draft-n-min 1 \
-ctk q8_0 -ctv q8_0 \
-ctkd q8_0 -ctvd q8_0 \
--temp 0.6 \
--reasoning on \
--reasoning-budget -1 \
--port 8080
For Mac with 96 GB RAM
# No MTP, compressed KV cache
./build/bin/llama-cli \
-m ~/Downloads/Hy3-IQ1_M.gguf \
-c 65536 \
-ngl 99 -fa on \
-ctk q8_0 -ctv q8_0 \
-p "Hello" -n 100 \
--temp 0.6
Flag reference
| Flag | What it does |
|---|---|
-m PATH |
Path to the GGUF model file |
-c N |
Context size in tokens. 65536 = 64K. Higher = more memory |
-ngl N |
Layers to offload to GPU. 99 = all layers on Metal |
-fa on |
Flash attention β reduces memory and speeds up attention |
-ctk q8_0 -ctv q8_0 |
KV cache in q8_0. Essential for 96 GB (saves ~20 GB) |
--temp N |
Sampling temperature. 0.0 = deterministic/greedy |
--reasoning on |
Enable thinking/reasoning (tag) |
--reasoning-budget N |
Max tokens for thinking. -1 = unlimited |
--spec-type draft-mtp |
MTP self-speculative decoding (*-mtp.gguf only) |
--spec-draft-n-max N |
Max draft tokens per MTP step |
Project structure
βββ 0001-add-hyv3-support.patch # Patch for 9 llama.cpp files (383 lines)
βββ src/models/hyv3.cpp # hy_v3 model implementation + MTP (388 lines)
βββ README.md # This file
Modified files in llama.cpp
| File | Change |
|---|---|
src/llama-arch.h |
New enum LLM_ARCH_HYV3 |
src/llama-arch.cpp |
Architecture name hy_v3 |
src/llama-model.cpp |
Model mapping + Neox rope type |
src/models/models.h |
llama_model_hyv3 class declaration |
src/models/hyv3.cpp |
New β load, forward, MTP draft head |
gguf-py/gguf/constants.py |
Arch enum + tensor list (28 hy_v3 tensors) |
gguf-py/gguf/tensor_mapping.py |
MTP tensor name mapping |
conversion/__init__.py |
HF β GGUF model name mapping |
common/chat.cpp |
Chat template parser (tool calls + reasoning) |
License
Apache 2.0. Same as the original Tencent/Hy3 model and AngelSlim's patches.
Long live open local AI. A 295B model running on a MacBook. π