--- base_model: Altworld/Astrea-R8-Chat-9B license: apache-2.0 library_name: gguf pipeline_tag: text-generation tags: - gguf - llama.cpp - qwen3.5 - chat - creative-writing - non-reasoning --- # Astrea R8 Chat 9B — Q8_0 GGUF This is an unofficial community Q8_0 GGUF conversion of [Altworld/Astrea-R8-Chat-9B](https://huggingface.co/Altworld/Astrea-R8-Chat-9B) for llama.cpp, with an explicit non-reasoning chat template. ## File | File | Quantization | Size | |---|---:|---:| | `Astrea-R8-Chat-9B-Q8_0.gguf` | Q8_0 | 9,527,501,280 bytes (8.87 GiB) | ## Embedded non-reasoning template The GGUF contains a hard non-reasoning Jinja template in `tokenizer.chat_template`; no external template file is required. Astrea's optional thinking mode was not reliable in local llama.cpp testing: simple prompts could consume hundreds of tokens before emitting ``, and often did not end reasoning at all, and simply responded as if reasoning was not enabled. The bundled template therefore always places a closed, empty thinking block in the prompt and does not expose an `enable_thinking` template variable. It also omits hidden reasoning when replaying assistant messages into conversation history. A standalone copy is included as `chat_template.jinja` for inspection. ## llama.cpp ```powershell llama-server.exe ` --model Astrea-R8-Chat-9B-Q8_0.gguf ` --jinja ` --reasoning off ` --reasoning-format none ` --ctx-size 32768 ` --n-gpu-layers all ` --temp 0.8 ` --top-p 1.0 ` --top-k 0 ` --min-p 0.025 ` --repeat-penalty 1.08 ``` The model metadata advertises a 262,144-token context window. Choose a context size appropriate for your available VRAM/RAM. The command above starts at a more conservative 32,768 tokens. I was able to easily run a much more ambitous setup with `-ngl all --fit off -c 147456 -np 4 --kv-unified` on a 16GB VRAM card (5070 Ti). ## Conversion notes 1. The original safetensors were converted to BF16 GGUF with llama.cpp's `convert_hf_to_gguf.py` using `--no-mtp`. The downloaded checkpoint did not contain the extra MTP-layer tensors declared by its configuration. 2. BF16 was quantized with `llama-quantize` using `Q8_0`. 3. llama.cpp's `gguf_new_metadata.py` embedded the hard non-reasoning template; this metadata-only copy did not requantize tensors. The tensor-only SHA-256 reported by `llama-gguf-hash` was identical before and after the metadata rewrite: ```text 20d213a0c5ee663ef6d02ffcff8d0b28cbff18b559c67aca5250cd5e6a22d624 ``` The final whole-file checksums are in `SHA256SUMS`. ## Validation The final GGUF was loaded directly by `llama-server` without `--chat-template-file`. Its exposed template matched the bundled standalone Jinja, and a request that explicitly supplied `enable_thinking=true` still returned normal content with no `reasoning_content`. ## License and attribution The source model is released under Apache-2.0. See `LICENSE` and `NOTICE`, and refer to the [source model card](https://huggingface.co/Altworld/Astrea-R8-Chat-9B) for its intended use, evaluation results, and limitations.