RobZombAI's picture
Upload 4 files
d3ff80a verified
|
Raw
History Blame Contribute Delete
8.08 kB
# Hy3 (hy_v3) for llama.cpp β€” macOS Metal support
Patches and binaries to run **Tencent Hy3 (hy_v3) 295B MoE** on **llama.cpp** with **Apple Silicon GPU (Metal)**.
---
## Why this project exists
llama.cpp supports dozens of architectures, but **Hy3 (hy_v3) wasn't one of them**. Tencent released Hy3 as open-source (Apache 2.0), AngelSlim quantized it to GGUF, but to run it on a Mac someone had to write the missing piece: hy_v3 architecture support in llama.cpp.
This repo contains **the patches that add hy_v3 to llama.cpp** β€” architecture detection, weight loading, MoE + shared expert forward pass, and MTP self-speculative decoding. All compiled with **Metal** for Apple Silicon GPU.
The goal is simple: **run a 295B model on a MacBook**. Not on the cloud, not on a cluster. On a laptop.
---
## Credits
This starts and ends with **AngelSlim** and their work on HuggingFace:
[**AngelSlim/Hy3-GGUF**](https://huggingface.co/AngelSlim/Hy3-GGUF) β€” GGUF-quantized model (IQ1_M and Q4_K_M), mixed-precision recipes, importance matrix, setup script, benchmarks, chat template. Without this work, this project wouldn't exist.
AngelSlim provided:
- The **base patches** for hy_v3 architecture in llama.cpp
- The **IQ1_M quantization** with mixed recipe (critical weights in Q8_0/Q6_K, experts in IQ1_M/IQ2_XXS)
- The **importance matrix** to allocate bits where they matter
- The **chat template** for tool calling and reasoning
This repo takes those patches, applies them to llama.cpp, and **builds them with Metal for macOS**.
**Thank you AngelSlim.** πŸ™Œ
---
## The model
| Detail | Value |
|--------|-------|
| **Architecture** | Hy3 (hy_v3) β€” Hunyuan V3 |
| **Developed by** | Tencent |
| **Parameters** | 295B |
| **Layers** | 81 (80 routed + 1 MTP) |
| **Experts** | 192 (8 active per token) |
| **Gating** | Sigmoid + correction bias + top-8 selection |
| **Quantization** | IQ1_M (AngelSlim mixed recipe) |
| **File size** | ~85 GB (with MTP) |
> Original HF model: [Tencent/Hy3](https://huggingface.co/tencent/Hy3)
> AngelSlim GGUF quant: [**AngelSlim/Hy3-GGUF**](https://huggingface.co/AngelSlim/Hy3-GGUF)
> Download: [Hy3-IQ1_M-mtp.gguf](https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M-mtp.gguf) (85 GB, IQ1_M with MTP)
**IQ1_M vs BF16 quality loss**: ~+0.3% PPL β€” **imperceptible**. Full benchmarks on [AngelSlim's HF page](https://huggingface.co/AngelSlim/Hy3-GGUF), file `assets/benchmark.png`.
---
## Why a MacBook?
| Mac | RAM | IQ1_M (85 GB) | MTP | Context | Notes |
|-----|-----|:---:|:---:|:--------:|-------|
| **M5 Max** | 128 GB | βœ… | βœ… | 64K | Everything on, comfortable |
| **M4 Max** | 128 GB | βœ… | βœ… | 64K | Everything on |
| **M3 Max** | 128 GB | βœ… | βœ… | 64K | Everything on |
| **MacBook Pro** | 128 GB | βœ… | βœ… | 64K | Runs on a laptop |
| **Mac Studio** | 96 GB | βœ… | ❌ | 64K | KV q8_0 only, no MTP |
| **MacBook Pro** | 96 GB | βœ… | ❌ | 64K | Same as above |
A **295B MoE running on a MacBook with 128 GB** is a concrete milestone for local AI:
- **No cloud**, no API keys, no subscriptions
- **Total privacy** β€” data never leaves your machine
- **No dedicated GPU needed** β€” Apple Silicon unified memory is enough
- **Portable** β€” no server rack, no cluster
With 96 GB it still works: the model is ~85 GB, leaving ~11 GB for the system. Just compress the KV cache (`-ctk q8_0 -ctv q8_0`) and skip MTP (which adds ~2 GB of weights plus a draft KV cache). With 128 GB everything runs β€” MTP included, with headroom.
---
## Download
### 1. The GGUF model
```bash
# IQ1_M with MTP (85 GB) β€” recommended for 128 GB
wget https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M-mtp.gguf
# IQ1_M without MTP (84 GB) β€” for 96 GB
wget https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M.gguf
```
### 2. The code
```bash
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
git checkout 19bba67c1
git apply /path/to/0001-add-hyv3-support.patch
cp /path/to/hyv3.cpp src/models/
```
Or clone the ready-made fork:
```bash
git clone https://github.com/RobZombAI/llama.cpp-metal_hyv3
cd llama.cpp-metal_hyv3
```
### 3. Build
```bash
mkdir build && cd build
cmake .. -DLLAMA_METAL=ON
make -j$(sysctl -n hw.logicalcpu)
```
---
## Commands
### CLI (base inference)
```bash
./build/bin/llama-cli \
-m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
-c 65536 \
-ngl 99 \
-fa on \
-ctk q8_0 -ctv q8_0 \
-p "Hello" \
-n 100 \
--temp 0.6
```
### CLI (with reasoning/thinking)
```bash
./build/bin/llama-cli \
-m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
-c 65536 \
-ngl 99 -fa on \
-ctk q8_0 -ctv q8_0 \
-p "Hello" \
-n 200 \
--temp 0.6 \
--reasoning on \
--reasoning-budget -1
```
### CLI (with MTP self-speculative β€” higher throughput)
```bash
./build/bin/llama-cli \
-m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
-c 65536 \
-ngl 99 -fa on \
--spec-type draft-mtp \
--spec-draft-n-max 3 \
--spec-draft-n-min 1 \
-ctk q8_0 -ctv q8_0 \
-ctkd q8_0 -ctvd q8_0 \
-p "Hello" \
-n 200 \
--temp 0.6
```
### Server (OpenAI-compatible API)
```bash
./build/bin/llama-server \
-m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
-c 65536 \
-ngl 99 -fa on \
-ctk q8_0 -ctv q8_0 \
--temp 0.6 \
--port 8080
```
Test API call:
```bash
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Hy3-IQ1_M-mtp",
"messages": [{"role": "user", "content": "Hello"}],
"temperature": 0.6,
"max_tokens": 200
}'
```
### Server (with reasoning + MTP)
```bash
./build/bin/llama-server \
-m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
-c 65536 \
-ngl 99 -fa on \
--spec-type draft-mtp \
--spec-draft-n-max 3 \
--spec-draft-n-min 1 \
-ctk q8_0 -ctv q8_0 \
-ctkd q8_0 -ctvd q8_0 \
--temp 0.6 \
--reasoning on \
--reasoning-budget -1 \
--port 8080
```
### For Mac with 96 GB RAM
```bash
# No MTP, compressed KV cache
./build/bin/llama-cli \
-m ~/Downloads/Hy3-IQ1_M.gguf \
-c 65536 \
-ngl 99 -fa on \
-ctk q8_0 -ctv q8_0 \
-p "Hello" -n 100 \
--temp 0.6
```
---
## Flag reference
| Flag | What it does |
|------|--------------|
| `-m PATH` | Path to the GGUF model file |
| `-c N` | Context size in tokens. `65536` = 64K. Higher = more memory |
| `-ngl N` | Layers to offload to GPU. `99` = all layers on Metal |
| `-fa on` | Flash attention β€” reduces memory and speeds up attention |
| `-ctk q8_0 -ctv q8_0` | KV cache in q8_0. **Essential for 96 GB** (saves ~20 GB) |
| `--temp N` | Sampling temperature. `0.0` = deterministic/greedy |
| `--reasoning on` | Enable thinking/reasoning (tag) |
| `--reasoning-budget N` | Max tokens for thinking. `-1` = unlimited |
| `--spec-type draft-mtp` | MTP self-speculative decoding (*-mtp.gguf only) |
| `--spec-draft-n-max N` | Max draft tokens per MTP step |
---
## Project structure
```
β”œβ”€β”€ 0001-add-hyv3-support.patch # Patch for 9 llama.cpp files (383 lines)
β”œβ”€β”€ src/models/hyv3.cpp # hy_v3 model implementation + MTP (388 lines)
└── README.md # This file
```
## Modified files in llama.cpp
| File | Change |
|------|--------|
| `src/llama-arch.h` | New enum `LLM_ARCH_HYV3` |
| `src/llama-arch.cpp` | Architecture name `hy_v3` |
| `src/llama-model.cpp` | Model mapping + Neox rope type |
| `src/models/models.h` | `llama_model_hyv3` class declaration |
| `src/models/hyv3.cpp` | **New** β€” load, forward, MTP draft head |
| `gguf-py/gguf/constants.py` | Arch enum + tensor list (28 hy_v3 tensors) |
| `gguf-py/gguf/tensor_mapping.py` | MTP tensor name mapping |
| `conversion/__init__.py` | HF β†’ GGUF model name mapping |
| `common/chat.cpp` | Chat template parser (tool calls + reasoning) |
## License
**Apache 2.0.** Same as the original [Tencent/Hy3](https://huggingface.co/tencent/Hy3) model and [AngelSlim](https://huggingface.co/AngelSlim)'s patches.
---
**Long live open local AI. A 295B model running on a MacBook. πŸŽ‰**