File size: 4,975 Bytes
7aeae89
 
 
210c173
7aeae89
 
 
 
 
 
 
 
 
210c173
7aeae89
210c173
7aeae89
 
 
 
 
 
 
 
 
 
 
 
 
 
210c173
7aeae89
 
 
 
 
 
210c173
7aeae89
 
 
210c173
 
 
 
7aeae89
 
 
 
 
 
 
 
 
 
 
 
210c173
 
 
 
 
 
 
 
7aeae89
 
 
210c173
7aeae89
 
 
 
 
210c173
 
 
 
 
7aeae89
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
210c173
7aeae89
210c173
7aeae89
 
 
 
210c173
7aeae89
 
 
 
 
 
210c173
7aeae89
 
 
 
 
 
210c173
 
 
7aeae89
 
 
210c173
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
---
license: apache-2.0
license_link: https://huggingface.co/tencent/Hy3/blob/main/LICENSE
thumbnail: https://huggingface.co/AtomicChat/Hy3-GGUF/resolve/main/hero.png
base_model:
- tencent/Hy3
base_model_relation: quantized
quantized_by: AtomicChat
pipeline_tag: text-generation
library_name: gguf
tags:
- atomic-chat
- hy3
- tencent
- gguf
- llama.cpp
- imatrix
- quantized
---

<center>

<div style="display:flex; justify-content:center; align-items:center; gap:2%; max-width:560px; margin:0 auto;">
<a href="https://atomic.chat" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/Hy3-GGUF/resolve/main/pill_atomic_v3.png" alt="Atomic Chat" style="width:100%; height:auto; max-width:186px;"></a>
<a href="https://discord.gg/8wGSsvmg4V" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/Hy3-GGUF/resolve/main/pill_discord_v3.png" alt="Join Discord" style="width:100%; height:auto; max-width:184px;"></a>
<a href="https://github.com/AtomicBot-ai/Atomic-Chat" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/Hy3-GGUF/resolve/main/pill_github_v3.png" alt="GitHub" style="width:100%; height:auto; max-width:141px;"></a>
</div>

<br/>

<img src="https://huggingface.co/AtomicChat/Hy3-GGUF/resolve/main/hero.png" alt="Hy3" style="width:100%; max-width:100%; height:auto; margin-bottom:0.6em;"/>

<div style="display:flex; justify-content:center; gap:0.5em;">
<a href="https://huggingface.co/tencent/Hy3"><strong>Base model: tencent/Hy3</strong></a>
</div>
</center>

**Hy3**, self-quantized to GGUF by [Atomic Chat](https://atomic.chat). Built straight from Tencent's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.

## Highlights

- **298.8B parameters**: the weights this repo quantizes.
- **Context length**: 262,144 tokens (256K), as published by Tencent.
- **80 layers**: Mixture-of-Experts.
- **Full imatrix ladder**: every quant is calibrated with an importance matrix, published here alongside the quants.

> [!NOTE]
> These GGUFs are **self-quantized from the original weights**, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.

> [!IMPORTANT]
> Always pass `--jinja` so the **Hy3 chat template** is applied. Without it the model can emit malformed turns.

## Model Overview

| Property | Value |
|---|---|
| Base model | `tencent/Hy3` |
| Parameters | 298.8B |
| Layers | 80 |
| Experts | 192 routed (top-8) |
| Context length | 262,144 tokens (256K) |
| Vocabulary | 120,832 |
| Modalities | Text |
| Architecture | Mixture-of-Experts, 192 experts (top-8), 64 attention heads over 8 KV heads, `HYV3ForCausalLM` |
| This repo | GGUF quants (imatrix); the importance matrix is published here as `imatrix-atomic.gguf`. Quants: `IQ1_M`, `Q4_K_M` |

<img src="https://huggingface.co/AtomicChat/Hy3-GGUF/resolve/main/benchmark.png" alt="Hy3 benchmark scores" style="width:100%; max-width:900px;"/>

Scores are Tencent's published results for the base `tencent/Hy3`, not our own measurements. Quantization preserves the large majority of this; `Q4_K_M` and up stay close to full precision.

## Choosing a quant

| Quant | Size | Notes |
|---|---|---|
| `IQ1_M` | 91.8 GB | Last resort, only if nothing else fits. |
| **`Q4_K_M`** | 184.7 GB | **Recommended default. Best balance of size, speed and quality.** |

> [!TIP]
> Pick the largest file that fits your (V)RAM with room for context. `Q4_K_M` is the sweet spot for most setups; `Q6_K` or `Q8_0` for maximum fidelity.

## Get started

Run Hy3 locally with:

- **[Atomic Chat](https://atomic.chat):** the easiest path. Open the app, search `AtomicChat/Hy3-GGUF`, pick a quant, hit **Use this model**.
- **llama.cpp:** `llama-server -hf AtomicChat/Hy3-GGUF:Q4_K_M --jinja -c 8192`
- **Ollama:** `ollama run hf.co/AtomicChat/Hy3-GGUF:Q4_K_M`
- **LM Studio / Jan:** search the repo id, download any quant.

## Best practices

| Parameter | Value |
|---|---|
| temperature | 0.9 |
| top_p | 1.0 |
| top_k | -1 |

Tencent's recommended sampling configuration for `tencent/Hy3`.

## Run in llama.cpp

```bash
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --target llama-cli llama-server
```

```bash
./llama.cpp/build/bin/llama-server \
    -hf AtomicChat/Hy3-GGUF:Q4_K_M \
    --jinja -ngl 99 -c 8192 -fa on
```

## How these were made

1. Download `tencent/Hy3` (original weights).
2. Convert to f16 GGUF with [llama.cpp](https://github.com/ggml-org/llama.cpp).
3. Build an importance matrix over our calibration corpus, published here as `imatrix-atomic.gguf`.
4. Quantize the ladder with `--imatrix`.

## License

Original model by Tencent, released under the Apache 2.0 license. Full terms: [Apache 2.0](https://huggingface.co/tencent/Hy3/blob/main/LICENSE). Quantized by Atomic Chat.