File size: 10,311 Bytes
f94a26a
 
 
 
 
 
 
 
 
 
 
 
 
 
08a6713
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f94a26a
 
08a6713
 
 
f94a26a
08a6713
 
 
 
 
 
 
 
f60f9a9
08a6713
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f60f9a9
 
 
08a6713
f60f9a9
 
08a6713
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f60f9a9
08a6713
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
---
license: apache-2.0
base_model: Qwen/Qwen3.5-9B
pipeline_tag: image-text-to-text
library_name: transformers
language:
- en
tags:
- qwen3.5
- chat
- creative-writing
- altworld
---

# Altworld Astrea: The New Frontier of Creative Writing

**[Try Astrea →](https://chat.altworld.io)** · **[Get an API key](https://altworld.io/developer)** · **[Read the docs →](https://astrea.altworld.io/docs/)**

Today, we are introducing **Altworld Astrea** — our first openly released
model. Astrea is a compact 9B-parameter model built for one thing: writing
that feels written. It sounds human, it keeps the story straight, and it
beats bigger models at both. You say something, it answers back.

**Altworldbench — overall writing quality.** Elo from 794 blind matchups on
our release pool — both rivals are bigger models built for creative writing:

| | Elo |
|---|---:|
| **Astrea 9B** | **1114** |
| Rocinante-X 12B | 1097 |
| Wayfarer-2 12B | 1045 |

Astrea is available today on [chat.altworld.io](https://chat.altworld.io),
and the full model weights are right here under Apache-2.0, so anyone can
download it and run it locally. An OpenAI-compatible
[API](https://astrea.altworld.io/docs/) is open with simple prepaid pricing.

## What makes Astrea different

### It sounds like a person

Most AI prose gives itself away in the first paragraph: the stock phrases,
the "it wasn't X — it was Y", the em-dash pileups, every reply opening the
same way. We count those tells. Astrea scores **2.5** — roughly two and a
half tells slipping into every thousand words. Every other model we tested
leaves between 8 and 19 per thousand.

### It understands the scene, not just the sentence

Judged blind, story against story, Astrea wins on the things that make
writing actually work: characters whose motivations hold together, a detail
planted early paying off later, the tone you asked for instead of generic
fantasy voice. That depth is what puts it ahead of models three billion
parameters bigger.

### It stays grounded

If the lantern went out, it stays out. If your sword was taken, it doesn't
reappear in your hand two scenes later. That's grounding — remembering what's
true in the story — and creative models are notoriously bad at it. Astrea
keeps **73 of 100** scenes straight. The strongest 12B prose model manages
**44**.

### All of it at 9B

Every rival on our board is both bigger than Astrea *and* purpose-built for
creative writing. Astrea isn't beating generalists — it's beating the
specialists, with three billion fewer parameters. That also means it runs
faster, serves cheaper, and fits on hardware the bigger models can't.

## Altworldbench

We test every release on **Altworldbench**, our own creative-writing
benchmark. Same prompts and settings for every model, every story judged
blind.

**Head-to-head wins** — share of blind matchups Astrea wins:

| | wins |
|---|---:|
| vs Rocinante-X | 52% |
| vs Wayfarer-2 | 65% |

**AI-sounding writing** — lower is better:

| | score |
|---|---:|
| **Astrea 9B** | **2.5** |
| Rocinante-X | 8.1 |
| Wayfarer-2 | 8.3 |
| Muse | 8.8 |
| Violet Lotus | 12.8 |

**Keeps the story straight** — out of 100:

| | score |
|---|---:|
| **Astrea 9B** | **73** |
| Violet Lotus | 68 |
| Muse | 64 |
| Wayfarer-2 | 61 |
| Rocinante-X | 44 |

**Writes the length you ask for** — % of the time:

| | % |
|---|---:|
| **Astrea 9B** | **93** |
| Rocinante-X | 88 |
| Violet Lotus | 83 |
| Wayfarer-2 | 78 |
| Muse | 25 |

All rivals are open 12B creative-writing models.

## See for yourself

Two samples from our test set — unedited, exactly as Astrea wrote them.

> **SCENE:** The rescue party returns with empty stretchers and one boot.
> **ACTION:** I ask the search captain to tell the families myself.

> The captain sets the boot on the table. It is a heavy boot, the kind the
> welders wear, with the steel cap dented in and the laces frayed. The
> stretchers are empty now, returned to the hangar. The rescue party came
> back in four, and one of them did not return.
>
> The families are waiting in the corridor. They have not eaten in days.
> The ration scrip is worthless now that half the station is gone, and the
> spin-section has slowed to a crawl. The air smells of recycled carbon and
> old sweat. You do not speak of the lost. You do not speak of the dead.
>
> The captain turns to you. He does not speak. The boot is on the table. He
> steps back.
>
> You stand. You take the boot. You walk into the corridor. The families
> look at you. There is nothing to say. You set the boot on the floor and
> say, "Your children did not come back."

> **SCENE:** Your partner swore your cargo of walrus ivory to two different
> buyers.
> **ACTION:** I go to the lesser buyer first and offer him the truth and a
> discount.

> The lesser buyer squints at you through a greasy half-door in Uppsala's
> timber shed. The air smells of pine pitch and old fish.
>
> He does not lie. The other party is already buying that ivory. Three
> measures from Helsingjar, not yours. You tell him the truth, offer him a
> discount of two dirhams. He laughs, spits. "The other man has my seal. He
> paid me in hacksilver and a vow over the sacred waters. You offer scraps,
> and I must burn them for luck."
>
> He turns away. The floorboards are wet. Frost lies thin on Lake Mälaren
> beyond the sheds, and the ice does not hold for a barge. The cargo
> remains in the gloom: twenty tusks, the colour of a seal's eye, unsworn
> to you now.

## Availability

- **Talk to Astrea** — chat with it now at
  [chat.altworld.io](https://chat.altworld.io) — no account needed for your
  first turns.
- **Build with the API** — an OpenAI-compatible endpoint at
  `api.altworld.io` — create a key and load credits at
  [altworld.io/developer](https://altworld.io/developer), then
  [start in the docs](https://astrea.altworld.io/docs/).
- **Play Astrea** — Astrea powers the voices of
  [Altworld](https://altworld.io), our life simulator and story sandbox.
- **Download the weights** — this repository: Apache-2.0, BF16 safetensors.

## Build with the API

The Astrea API is **OpenAI-compatible**: if your code already talks to an
OpenAI-style endpoint, point it at our base URL, swap the model name, and it
works — streaming included.

```bash
curl https://api.altworld.io/v1/chat/completions \
  -H "Authorization: Bearer $ASTREA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "astrea",
    "messages": [{"role": "user", "content": "Open a story in a lighthouse, first person, no cliches."}],
    "stream": true
  }'
```

## Model details

Altworld Astrea (release name Astrea-R8-Chat-9B) was trained on top of the
Qwen3.5-9B base model using multi-round LoRA supervised fine-tuning,
alternating conversational and narrative rounds and checking every round
against Altworldbench. The released checkpoint is the merged BF16 model.

The chat template works bare: no system prompt, no flags, it just answers —
with no visible reasoning block. Explicit thinking is available per-request
with `enable_thinking=True`.

Recommended sampling:

```json
"temperature": 0.8,
"min_p": 0.025,
"repetition_penalty": 1.08
```

Use `temperature 0.2` for factual chat.

## Limitations

- Like any compact model, Astrea can miss nuance or state things confidently
  and wrongly. Verify anything high-stakes independently.
- English dominates its training and evaluation. Other languages will work,
  but with weaker guarantees.
- This release is text-first: the architecture inherits vision support, but
  visual behavior was not part of the release evaluation.
- Astrea is not certified for medical, legal, financial, or other
  high-stakes use.

We plan to keep improving and openly releasing Astrea models — tell us what
to fix. The fastest way to reach us is the Altworld Discord, or open a
Discussion on this repository.

---

## The fine print

**Evaluation protocol.** The pooled Elo is 1113.7 with a 95% bootstrap
interval of 1089.1–1140.9 on a frozen seven-model pool (794 comparisons for
Astrea's entry). The interval overlaps Astrea R8 and Rocinante-X-12B-v1, so
the supported claim is #1 on point estimate and statistical parity with the
top cluster — not a statistically separated win. Pairwise judgments were
made blind by GLM-5p1 and DeepSeek-V4-Flash, with Kimi-K2.6 adjudicating
disagreements; grounding, slop, length/POV/tense, and bare-chat checks
combine deterministic rules with LLM judging. Separately from the narrative
pool: 92.27% pass on 207 held-out bare-chat prompts, 0.00% accidental
fiction, 20/20 on unseen identity paraphrases. One constructed multi-turn
memory probe answered indirectly at temperature 0.2 while fine at 0.8 —
test your own dialogue distributions.

**Running it locally.** Validated in production on vLLM 0.25 (A40, and
2×RTX 3090 with `--tensor-parallel-size 2` — set `NCCL_P2P_DISABLE=1` on
consumer dual-GPU boards):

```bash
vllm serve Altworld/Astrea-R8-Chat-9B \
  --dtype bfloat16 --max-model-len 32768 --gpu-memory-utilization 0.90
```

The base architecture natively supports 262k context; the hybrid
linear-attention layers keep the KV cache small, so serve what your VRAM
affords. With Transformers, load via `AutoProcessor` +
`AutoModelForMultimodalLM` and apply the bundled chat template
(`enable_thinking` defaults off). An MLX-LM smoke run behaved worse than
vLLM at temperature 0.2 despite numerically matched logits — validate your
exact engine. GGUF quants are planned; llama.cpp supports the Qwen3.5
family.

**Lineage and license.** Base model
[Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), Apache-2.0; this
derivative is released under the same license — see `NOTICE`. Fine-tuning
preference data used gold responses from MIT-licensed open-weight models
(GLM-5.2, DeepSeek-V4-Flash) with on-policy rejected sides; no closed-model
outputs anywhere. Early-stage data ancestry credits: Athanorlite-DPO
(CC-BY-4.0), jondurbin/gutenberg-dpo-v0.1 (CC-BY-4.0),
nbeerbower/gutenberg2-dpo and gutenberg-moderne-dpo (Apache-2.0),
Pageshift-Entertainment/LongPage, allenai/tulu-3-sft-mixture (ODC-BY,
license-filtered), nothingiisreal/Reddit-Dirty-And-WritingPrompts
(Apache-2.0, SFW file only). Thanks to the Qwen team for the base model.