Michionlion's picture
Add initial release
e14e0a2 verified
|
Raw
History Blame Contribute Delete
3.1 kB
---
base_model: Altworld/Astrea-R8-Chat-9B
license: apache-2.0
library_name: gguf
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- qwen3.5
- chat
- creative-writing
- non-reasoning
---
# Astrea R8 Chat 9B — Q8_0 GGUF
This is an unofficial community Q8_0 GGUF conversion of
[Altworld/Astrea-R8-Chat-9B](https://huggingface.co/Altworld/Astrea-R8-Chat-9B)
for llama.cpp, with an explicit non-reasoning chat template.
## File
| File | Quantization | Size |
|---|---:|---:|
| `Astrea-R8-Chat-9B-Q8_0.gguf` | Q8_0 | 9,527,501,280 bytes (8.87 GiB) |
## Embedded non-reasoning template
The GGUF contains a hard non-reasoning Jinja template in
`tokenizer.chat_template`; no external template file is required.
Astrea's optional thinking mode was not reliable in local llama.cpp testing:
simple prompts could consume hundreds of tokens before emitting `</think>`, and often did not end reasoning at all, and simply responded as if reasoning was not enabled.
The bundled template therefore always places a closed, empty thinking block in
the prompt and does not expose an `enable_thinking` template variable. It also
omits hidden reasoning when replaying assistant messages into conversation
history. A standalone copy is included as `chat_template.jinja` for inspection.
## llama.cpp
```powershell
llama-server.exe `
--model Astrea-R8-Chat-9B-Q8_0.gguf `
--jinja `
--reasoning off `
--reasoning-format none `
--ctx-size 32768 `
--n-gpu-layers all `
--temp 0.8 `
--top-p 1.0 `
--top-k 0 `
--min-p 0.025 `
--repeat-penalty 1.08
```
The model metadata advertises a 262,144-token context window. Choose a context
size appropriate for your available VRAM/RAM. The command above starts at a
more conservative 32,768 tokens. I was able to easily run a much more ambitous setup with `-ngl all --fit off -c 147456 -np 4 --kv-unified` on a 16GB VRAM card (5070 Ti).
## Conversion notes
1. The original safetensors were converted to BF16 GGUF with llama.cpp's
`convert_hf_to_gguf.py` using `--no-mtp`. The downloaded checkpoint did not
contain the extra MTP-layer tensors declared by its configuration.
2. BF16 was quantized with `llama-quantize` using `Q8_0`.
3. llama.cpp's `gguf_new_metadata.py` embedded the hard non-reasoning template;
this metadata-only copy did not requantize tensors.
The tensor-only SHA-256 reported by `llama-gguf-hash` was identical before and
after the metadata rewrite:
```text
20d213a0c5ee663ef6d02ffcff8d0b28cbff18b559c67aca5250cd5e6a22d624
```
The final whole-file checksums are in `SHA256SUMS`.
## Validation
The final GGUF was loaded directly by `llama-server` without
`--chat-template-file`. Its exposed template matched the bundled standalone
Jinja, and a request that explicitly supplied `enable_thinking=true` still
returned normal content with no `reasoning_content`.
## License and attribution
The source model is released under Apache-2.0. See `LICENSE` and `NOTICE`, and
refer to the [source model card](https://huggingface.co/Altworld/Astrea-R8-Chat-9B)
for its intended use, evaluation results, and limitations.