RobZombAI's picture
Upload 4 files
d3ff80a verified
|
Raw
History Blame Contribute Delete
4.1 kB
metadata
license: apache-2.0
language:
  - en
  - zh
pipeline_tag: text-generation
tags:
  - moe
  - hy3
  - hy-v3
  - huggingface
  - llama.cpp
  - apple-silicon
  - metal
  - local-ai
  - 295b
base_model: tencent/Hy3
datasets:
  - AngelSlim/Hy3-GGUF

Hy3 (hy_v3) β€” macOS Metal Build

Run Tencent Hy3 295B MoE on a MacBook via llama.cpp + Metal.

Not in the cloud. Not on a cluster. On a laptop.

πŸ”₯ The Numbers

What Value
Architecture Hy3 (hy_v3) β€” Hunyuan V3
Parameters 295B
Layers 81 (80 routed + 1 MTP)
Experts 192 (8 active per token)
Quantization IQ1_M (mixed recipe)
File size ~85 GB
RAM needed 96-128 GB (Apple Silicon)
License Apache 2.0

πŸ† Why This Matters

A 295B MoE model running on a single MacBook is a milestone for open local AI:

  • βœ… No cloud, no API keys, no subscriptions
  • βœ… Total privacy β€” data never leaves your machine
  • βœ… No dedicated GPU β€” Apple Silicon unified memory is enough
  • βœ… Portable β€” runs on a laptop, not a server rack

πŸ“₯ Download the GGUF

The quantized model comes from AngelSlim/Hy3-GGUF:

# IQ1_M with MTP (85 GB) β€” recommended for 128 GB
wget https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M-mtp.gguf

# IQ1_M without MTP (84 GB) β€” for 96 GB
wget https://huggingface.co/AngelSlim/Hy3-GGUF/resolve/main/Hy3-IQ1_M.gguf

πŸ› οΈ Build & Run

1. Clone the patched llama.cpp

git clone https://github.com/RobZombAI/llama.cpp-metal_hyv3
cd llama.cpp-metal_hyv3

2. Build with Metal

mkdir build && cd build
cmake .. -DLLAMA_METAL=ON
make -j$(sysctl -n hw.logicalcpu)

3. Run inference

./build/bin/llama-cli \
    -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
    -c 65536 \
    -ngl 99 \
    -fa on \
    -ctk q8_0 -ctv q8_0 \
    -p "Hello" \
    -n 100 \
    --temp 0.6

With reasoning enabled

./build/bin/llama-cli \
    -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
    -c 65536 -ngl 99 -fa on \
    -ctk q8_0 -ctv q8_0 \
    -p "Explain quantum computing" \
    -n 300 --temp 0.6 \
    --reasoning on --reasoning-budget -1

With MTP self-speculative decoding (higher throughput)

./build/bin/llama-cli \
    -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
    -c 65536 -ngl 99 -fa on \
    --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-n-min 1 \
    -ctk q8_0 -ctv q8_0 \
    -ctkd q8_0 -ctvd q8_0 \
    -p "Hello" -n 200 --temp 0.6

OpenAI-compatible server

./build/bin/llama-server \
    -m ~/Downloads/Hy3-IQ1_M-mtp.gguf \
    -c 65536 -ngl 99 -fa on \
    -ctk q8_0 -ctv q8_0 \
    --temp 0.6 --port 8080

For 96 GB Macs (no MTP, compressed KV)

./build/bin/llama-cli \
    -m ~/Downloads/Hy3-IQ1_M.gguf \
    -c 65536 -ngl 99 -fa on \
    -ctk q8_0 -ctv q8_0 \
    -p "Hello" -n 100 --temp 0.6

πŸ“Š Hardware Requirements

Mac RAM MTP Context Notes
M5/M4/M3 Max 128 GB βœ… 64K Everything on
MacBook Pro 128 GB βœ… 64K Runs on a laptop
Mac Studio 96 GB ❌ 64K KV q8_0, no MTP
MacBook Pro 96 GB ❌ 64K Same

πŸ™ Credits

  • Tencent β€” original Hy3 model (Apache 2.0)
  • AngelSlim β€” GGUF quantization, mixed recipes, importance matrix, base llama.cpp patches, benchmarks. The real MVP.
  • RobZombAI β€” macOS Metal build, patches integration, testing

πŸ“š Resources

πŸ“œ License

Apache 2.0 β€” same as the original Tencent Hy3 model and AngelSlim's patches.


Long live open local AI. πŸŽ‰