File size: 2,908 Bytes
5f3b0aa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
---
license: apache-2.0
base_model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
library_name: llama.cpp
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- rocm
- rocmfpx
- amd
- strix-halo
- qwen3.8
- qwen35
- imatrix
- thinking
- uncensored
---

# Qwen3.8-27B AEON ULTIMATE — ROCmFPX iMatrix GGUF

ROCmFPX iMatrix quantizations of the full BF16 checkpoint published by
[Aeon / AEON-7](https://huggingface.co/AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16).

## Attribution

The source weights are Aeon's `Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16`,
pinned at Hub revision
`8f76e82ed7ef4de7735f5d4148fce7b643b00fae`. Aeon deserves attribution for the
BF16 model and its model work. This repository contains derived GGUF
quantizations produced by vmlinux with the ROCmFPX toolchain; it is not a new
training run or a claim of ownership of the source model.

The source model declares Apache 2.0 licensing. Review the source model card
and applicable terms before redistribution or deployment.

## Quantizations

| File | Preset | Size | SHA-256 |
| --- | --- | ---: | --- |
| `Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP4-iMatrix.gguf` | `Q4_0_ROCMFP4` | 17,735,469,440 bytes | `34c04aaec2399fab8179a921d0ea1e9d3b93da213201ab0ec2930ee1420307b0` |
| `Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP6-iMatrix.gguf` | `Q6_0_ROCMFPX` | 22,528,383,360 bytes | `d096fdd0ece150cefa793a8fdae55076b63d61d9e135571de6c73b016a641ba1` |
| `Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP8-iMatrix.gguf` | `Q8_0_ROCMFPX` | 28,193,397,120 bytes | `10eff5653c51e23a8a816761feea69b235c9003e80821759ecef41814843a761` |

All three use the same model-specific importance matrix:
`Qwen3.8-27B-AEON-ULTIMATE-iMatrix.imatrix.gguf`. The matrix was generated
from 339 chunks of 512 tokens using the shared calibration corpus, and each
quantizer consumed 496 entries.

## Runtime

These are experimental ROCmFPX tensor types and require a compatible
ROCmFPX-enabled llama.cpp build. Stock upstream llama.cpp will not load them.

Example ROCm0 invocation:

```bash
hf download vmlinux/Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX-GGUF \
  Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP4-iMatrix.gguf \
  --local-dir ./Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX

./llama-completion \
  -m ./Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP4-iMatrix.gguf \
  -dev ROCm0 -ngl all -c 8192 -n 256 -p "Hello."
```

The GGUFs retain the native one-layer MTP head. Thinking is enabled by the
embedded Qwen template by default; pass the appropriate chat-template kwargs
when an application needs thinking disabled.

## Validation and provenance

All three files loaded and generated a short completion on ROCm0 with all
layers offloaded. Detailed public build information is in
[`BUILD_RESULTS.md`](BUILD_RESULTS.md), with exact hashes in
[`SHA256SUMS`](SHA256SUMS) and source/toolchain details in
[`PROVENANCE.md`](PROVENANCE.md).