Michionlion's picture
Add initial release
e14e0a2 verified
|
Raw
History Blame Contribute Delete
3.1 kB
metadata
base_model: Altworld/Astrea-R8-Chat-9B
license: apache-2.0
library_name: gguf
pipeline_tag: text-generation
tags:
  - gguf
  - llama.cpp
  - qwen3.5
  - chat
  - creative-writing
  - non-reasoning

Astrea R8 Chat 9B — Q8_0 GGUF

This is an unofficial community Q8_0 GGUF conversion of Altworld/Astrea-R8-Chat-9B for llama.cpp, with an explicit non-reasoning chat template.

File

File Quantization Size
Astrea-R8-Chat-9B-Q8_0.gguf Q8_0 9,527,501,280 bytes (8.87 GiB)

Embedded non-reasoning template

The GGUF contains a hard non-reasoning Jinja template in tokenizer.chat_template; no external template file is required.

Astrea's optional thinking mode was not reliable in local llama.cpp testing: simple prompts could consume hundreds of tokens before emitting </think>, and often did not end reasoning at all, and simply responded as if reasoning was not enabled. The bundled template therefore always places a closed, empty thinking block in the prompt and does not expose an enable_thinking template variable. It also omits hidden reasoning when replaying assistant messages into conversation history. A standalone copy is included as chat_template.jinja for inspection.

llama.cpp

llama-server.exe `
  --model Astrea-R8-Chat-9B-Q8_0.gguf `
  --jinja `
  --reasoning off `
  --reasoning-format none `
  --ctx-size 32768 `
  --n-gpu-layers all `
  --temp 0.8 `
  --top-p 1.0 `
  --top-k 0 `
  --min-p 0.025 `
  --repeat-penalty 1.08

The model metadata advertises a 262,144-token context window. Choose a context size appropriate for your available VRAM/RAM. The command above starts at a more conservative 32,768 tokens. I was able to easily run a much more ambitous setup with -ngl all --fit off -c 147456 -np 4 --kv-unified on a 16GB VRAM card (5070 Ti).

Conversion notes

  1. The original safetensors were converted to BF16 GGUF with llama.cpp's convert_hf_to_gguf.py using --no-mtp. The downloaded checkpoint did not contain the extra MTP-layer tensors declared by its configuration.
  2. BF16 was quantized with llama-quantize using Q8_0.
  3. llama.cpp's gguf_new_metadata.py embedded the hard non-reasoning template; this metadata-only copy did not requantize tensors.

The tensor-only SHA-256 reported by llama-gguf-hash was identical before and after the metadata rewrite:

20d213a0c5ee663ef6d02ffcff8d0b28cbff18b559c67aca5250cd5e6a22d624

The final whole-file checksums are in SHA256SUMS.

Validation

The final GGUF was loaded directly by llama-server without --chat-template-file. Its exposed template matched the bundled standalone Jinja, and a request that explicitly supplied enable_thinking=true still returned normal content with no reasoning_content.

License and attribution

The source model is released under Apache-2.0. See LICENSE and NOTICE, and refer to the source model card for its intended use, evaluation results, and limitations.