| Hy3 (hy_v3) for llama.cpp β macOS Metal support |
| Patches and binaries to run Tencent Hy3 (hy_v3) 295B MoE on llama.cpp with Apple Silicon GPU (Metal). |
|
|
| Why this project exists |
| llama.cpp supports dozens of architectures, but Hy3 (hy_v3) wasn't one of them. Tencent released Hy3 as open-source (Apache 2.0), AngelSlim quantized it to GGUF, but to run it on a Mac someone had to write the missing piece: hy_v3 architecture support in llama.cpp. |
|
|
| This repo contains the patches that add hy_v3 to llama.cpp β architecture detection, weight loading, MoE + shared expert forward pass, and MTP self-speculative decoding. All compiled with Metal for Apple Silicon GPU. |
| |
| The goal is simple: run a 295B model on a MacBook. Not on the cloud, not on a cluster. On a laptop. |
| |
| Credits |
| This starts and ends with AngelSlim and their work on HuggingFace: |
| |
| AngelSlim/Hy3-GGUF β GGUF-quantized model (IQ1_M and Q4_K_M), mixed-precision recipes, importance matrix, setup script, benchmarks, chat template. Without this work, this project wouldn't exist. |
|
|
| AngelSlim provided: |
|
|
| The base patches for hy_v3 architecture in llama.cpp |
| The IQ1_M quantization with mixed recipe (critical weights in Q8_0/Q6_K, experts in IQ1_M/IQ2_XXS) |
| The importance matrix to allocate bits where they matter |
| The chat template for tool calling and reasoning |
| This repo takes those patches, applies them to llama.cpp, and builds them with Metal for macOS. |
|
|
| Thank you AngelSlim. π |
|
|
| The model |
| Detail Value |
| Architecture Hy3 (hy_v3) β Hunyuan V3 |
| Developed by Tencent |
| Parameters 295B |
| Layers 81 (80 routed + 1 MTP) |
| Experts 192 (8 active per token) |
| Gating Sigmoid + correction bias + top-8 selection |
| Quantization IQ1_M (AngelSlim mixed recipe) |
| File size ~85 GB (with MTP) |
| Original HF model: Tencent/Hy3 AngelSlim GGUF quant: AngelSlim/Hy3-GGUF Download: Hy3-IQ1_M-mtp.gguf (85 GB, IQ1_M with MTP) |
|
|
| IQ1_M vs BF16 quality loss: ~+0.3% PPL β imperceptible. Full benchmarks on AngelSlim's HF page, file assets/benchmark.png. |
| |
| Why a MacBook? |
| Mac RAM IQ1_M (85 GB) MTP Context Notes |
| M5 Max 128 GB β
β
64K Everything on, comfortable |
| M4 Max 128 GB β
β
64K Everything on |
| M3 Max 128 GB β
β
64K Everything on |
| MacBook Pro 128 GB β
β
64K Runs on a laptop |
| Mac Studio 96 GB β
β 64K KV q8_0 only, no MTP |
| MacBook Pro 96 GB β
β 64K Same as above |
| A 295B MoE running on a MacBook with 128 GB is a concrete milestone for local AI: |
| |
| No cloud, no API keys, no subscriptions |
| Total privacy β data never leaves your machine |
| No dedicated GPU needed β Apple Silicon unified memory is enough |
| Portable β no server rack, no cluster |
| With 96 GB it still works: the model is ~85 GB, leaving ~11 GB for the system. Just compress the KV cache (-ctk q8_0 -ctv q8_0) and skip MTP (which adds ~2 GB of weights plus a draft KV cache). With 128 GB everything runs β MTP included, with headroom. |
| |
| Download |
| 1. The GGUF model |
| # IQ1_M with MTP (85 GB) β recommended for 128 GB |
| wget https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M-mtp.gguf |
| |
| # IQ1_M without MTP (84 GB) β for 96 GB |
| wget https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M.gguf |
| 2. The code |
| git clone https://github.com/ggml-org/llama.cpp |
| cd llama.cpp |
| git checkout 19bba67c1 |
| git apply /path/to/0001-add-hyv3-support.patch |
| cp /path/to/hyv3.cpp src/models/ |
| Or clone the ready-made fork: |
| |
| git clone https://github.com/<your-username>/llama.cpp-hyv3 |
| cd llama.cpp-hyv3 |
| 3. Build |
| mkdir build && cd build |
| cmake .. -DLLAMA_METAL=ON |
| make -j$(sysctl -n hw.logicalcpu) |
| Commands |
| CLI (base inference) |
| ./build/bin/llama-cli \ |
| -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \ |
| -c 65536 \ |
| -ngl 99 \ |
| -fa on \ |
| -ctk q8_0 -ctv q8_0 \ |
| -p "Hello" \ |
| -n 100 \ |
| --temp 0.6 |
| CLI (with reasoning/thinking) |
| ./build/bin/llama-cli \ |
| -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \ |
| -c 65536 \ |
| -ngl 99 -fa on \ |
| -ctk q8_0 -ctv q8_0 \ |
| -p "Hello" \ |
| -n 200 \ |
| --temp 0.6 \ |
| --reasoning on \ |
| --reasoning-budget -1 |
| |
| Server (OpenAI-compatible API) |
| ./build/bin/llama-server \ |
| -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \ |
| -c 65536 \ |
| -ngl 99 -fa on \ |
| -ctk q8_0 -ctv q8_0 \ |
| --temp 0.6 \ |
| --port 8080 |
| Test API call: |
| |
| curl http://localhost:8080/v1/chat/completions \ |
| -H "Content-Type: application/json" \ |
| -d '{ |
| "model": "Hy3-IQ1_M-mtp", |
| "messages": [{"role": "user", "content": "Hello"}], |
| "temperature": 0.6, |
| "max_tokens": 200 |
| }' |
| |
| For Mac with 96 GB RAM |
| # No MTP, compressed KV cache |
| ./build/bin/llama-cli \ |
| -m ~/Downloads/Hy3-IQ1_M.gguf \ |
| -c 65536 \ |
| -ngl 99 -fa on \ |
| -ctk q8_0 -ctv q8_0 \ |
| -p "Hello" -n 100 \ |
| --temp 0.6 |
| Flag reference |
| Flag What it does |
| -m PATH Path to the GGUF model file |
| -c N Context size in tokens. 65536 = 64K. Higher = more memory |
| -ngl N Layers to offload to GPU. 99 = all layers on Metal |
| -fa on Flash attention β reduces memory and speeds up attention |
| -ctk q8_0 -ctv q8_0 KV cache in q8_0. Essential for 96 GB (saves ~20 GB) |
| --temp N Sampling temperature. 0.0 = deterministic/greedy |
| --reasoning on Enable thinking/reasoning (tag) |
| --reasoning-budget N Max tokens for thinking. -1 = unlimited |
| --spec-type draft-mtp MTP self-speculative decoding (*-mtp.gguf only) |
| --spec-draft-n-max N Max draft tokens per MTP step |
| Project structure |
| βββ 0001-add-hyv3-support.patch # Patch for 9 llama.cpp files (383 lines) |
| βββ src/models/hyv3.cpp # hy_v3 model implementation + MTP (388 lines) |
| βββ README.md # This file |
| Modified files in llama.cpp |
| File Change |
| src/llama-arch.h New enum LLM_ARCH_HYV3 |
| src/llama-arch.cpp Architecture name hy_v3 |
| src/llama-model.cpp Model mapping + Neox rope type |
| src/models/models.h llama_model_hyv3 class declaration |
| src/models/hyv3.cpp New β load, forward, MTP draft head |
| gguf-py/gguf/constants.py Arch enum + tensor list (28 hy_v3 tensors) |
| gguf-py/gguf/tensor_mapping.py MTP tensor name mapping |
| conversion/__init__.py HF β GGUF model name mapping |
| common/chat.cpp Chat template parser (tool calls + reasoning) |
| License |
| Apache 2.0. Same as the original Tencent/Hy3 model and AngelSlim's patches. |
| |
| Long live open local AI. A 295B model running on a MacBook. π |
|
|