File size: 2,908 Bytes
5f3b0aa | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 | ---
license: apache-2.0
base_model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
library_name: llama.cpp
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- rocm
- rocmfpx
- amd
- strix-halo
- qwen3.8
- qwen35
- imatrix
- thinking
- uncensored
---
# Qwen3.8-27B AEON ULTIMATE — ROCmFPX iMatrix GGUF
ROCmFPX iMatrix quantizations of the full BF16 checkpoint published by
[Aeon / AEON-7](https://huggingface.co/AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16).
## Attribution
The source weights are Aeon's `Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16`,
pinned at Hub revision
`8f76e82ed7ef4de7735f5d4148fce7b643b00fae`. Aeon deserves attribution for the
BF16 model and its model work. This repository contains derived GGUF
quantizations produced by vmlinux with the ROCmFPX toolchain; it is not a new
training run or a claim of ownership of the source model.
The source model declares Apache 2.0 licensing. Review the source model card
and applicable terms before redistribution or deployment.
## Quantizations
| File | Preset | Size | SHA-256 |
| --- | --- | ---: | --- |
| `Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP4-iMatrix.gguf` | `Q4_0_ROCMFP4` | 17,735,469,440 bytes | `34c04aaec2399fab8179a921d0ea1e9d3b93da213201ab0ec2930ee1420307b0` |
| `Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP6-iMatrix.gguf` | `Q6_0_ROCMFPX` | 22,528,383,360 bytes | `d096fdd0ece150cefa793a8fdae55076b63d61d9e135571de6c73b016a641ba1` |
| `Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP8-iMatrix.gguf` | `Q8_0_ROCMFPX` | 28,193,397,120 bytes | `10eff5653c51e23a8a816761feea69b235c9003e80821759ecef41814843a761` |
All three use the same model-specific importance matrix:
`Qwen3.8-27B-AEON-ULTIMATE-iMatrix.imatrix.gguf`. The matrix was generated
from 339 chunks of 512 tokens using the shared calibration corpus, and each
quantizer consumed 496 entries.
## Runtime
These are experimental ROCmFPX tensor types and require a compatible
ROCmFPX-enabled llama.cpp build. Stock upstream llama.cpp will not load them.
Example ROCm0 invocation:
```bash
hf download vmlinux/Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX-GGUF \
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP4-iMatrix.gguf \
--local-dir ./Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX
./llama-completion \
-m ./Qwen3.8-27B-AEON-ULTIMATE-ROCmFPX/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-ROCmFP4-iMatrix.gguf \
-dev ROCm0 -ngl all -c 8192 -n 256 -p "Hello."
```
The GGUFs retain the native one-layer MTP head. Thinking is enabled by the
embedded Qwen template by default; pass the appropriate chat-template kwargs
when an application needs thinking disabled.
## Validation and provenance
All three files loaded and generated a short completion on ROCm0 with all
layers offloaded. Detailed public build information is in
[`BUILD_RESULTS.md`](BUILD_RESULTS.md), with exact hashes in
[`SHA256SUMS`](SHA256SUMS) and source/toolchain details in
[`PROVENANCE.md`](PROVENANCE.md).
|