File size: 3,432 Bytes
4261616
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
---
license: apache-2.0
base_model: Altworld/Astrea-R8-Chat-9B
pipeline_tag: text-generation
language:
- en
tags:
- gguf
- llama.cpp
- qwen3.5
- chat
- creative-writing
- altworld
---

# Astrea R8 Chat 9B — GGUF

Official text-only GGUF builds of
[Altworld/Astrea-R8-Chat-9B](https://huggingface.co/Altworld/Astrea-R8-Chat-9B),
a compact conversational and creative-writing model built on Qwen3.5-9B.

**[Try Astrea](https://chat.altworld.io)** ·
**[API](https://altworld.io/developer)** ·
**[Documentation](https://astrea.altworld.io/docs/)**

## Files

| Quant | Size | Use case |
|---|---:|---|
| **Q5_K_M** | 6.02 GiB | Recommended balance of quality and size |
| Q4_K_M | 5.24 GiB | Smaller and faster; passed the same chat-routing smoke tests |
| Q6_K | 6.85 GiB | Higher fidelity |
| Q8_0 | 8.87 GiB | Highest-fidelity quant in this repository |

These are single-file, text-only models. The inherited vision components are
not included because Astrea's visual behavior was not part of the release
evaluation.

## Run with llama.cpp

Use a current build of
[llama.cpp](https://github.com/ggml-org/llama.cpp); Qwen3.5 support is recent.

```bash
llama-cli \
  -hf Altworld/Astrea-R8-Chat-9B-GGUF:Q5_K_M \
  -cnv \
  --reasoning off \
  --temp 0.8 \
  --min-p 0.025 \
  --repeat-penalty 1.08
```

For an OpenAI-compatible local server:

```bash
llama-server \
  -hf Altworld/Astrea-R8-Chat-9B-GGUF:Q5_K_M \
  --reasoning off \
  --temp 0.8 \
  --min-p 0.025 \
  --repeat-penalty 1.08
```

Use `--temp 0.2` for factual chat. The model carries its official chat
template inside the GGUF and does not require a system prompt. Reasoning is
disabled in the examples because Astrea's release behavior is direct response
without a visible thinking block.

## Validation

Each file was loaded and generated from with llama.cpp. The checks covered:

- bare greeting with no system prompt;
- native Astrea/Altworld identity;
- ordinary factual chat;
- supportive conversation;
- creative-writing routing from a plain request with no scene card;
- multi-turn recall using the embedded chat template.

Q4_K_M, the most aggressive quant here, correctly answered `Good afternoon.`
as normal chat and produced a coherent story from `Write a short story about a
dragon who collects spoons.` Q5_K_M correctly recalled a name across turns.

The GGUFs declare the base architecture's 262,144-token context. Actual usable
context depends on your hardware, llama.cpp build, and runtime settings.

## Conversion details

Converted from the official merged BF16 release with llama.cpp commit
[`76f46ad`](https://github.com/ggml-org/llama.cpp/commit/76f46ad29d61fd8c1401e8221842934bf62a6064).
The source checkpoint contains no MTP tensors, so conversion used `--no-mtp`.
Quantization used llama.cpp's `Q4_K_M`, `Q5_K_M`, `Q6_K`, and `Q8_0` presets.

## About Astrea

Astrea was trained on top of Qwen3.5-9B using multi-round LoRA supervised
fine-tuning, alternating conversational and narrative rounds and checking each
round against Altworldbench. See the
[BF16 model card](https://huggingface.co/Altworld/Astrea-R8-Chat-9B) for
benchmarks, methodology, examples, limitations, and the evaluation protocol.

Astrea is English-first and can still hallucinate. Verify high-stakes claims
independently. Engine and quantization differences can affect output, so test
your own prompts and settings.

## License

Apache-2.0. See `LICENSE` and `NOTICE`.