Instructions to use julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF # Run inference directly in the terminal: llama cli -hf julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF # Run inference directly in the terminal: llama cli -hf julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF # Run inference directly in the terminal: ./llama-cli -hf julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
Use Docker
docker model run hf.co/julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
- LM Studio
- Jan
- vLLM
How to use julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
- Ollama
How to use julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF with Ollama:
ollama run hf.co/julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
- Unsloth Studio
How to use julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF to start chatting
- Pi
How to use julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF with Docker Model Runner:
docker model run hf.co/julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
- Lemonade
How to use julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
Run and chat with the model
lemonade run user.Qwen-3.8-27B-ROCmFP4-FAST-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
Run Hermes
hermes
- Atomic Chat
Qwen 3.8 27B ROCmFP4_FAST (GGUF for AMD Strix Halo)
This repository contains the optimized ROCmFP4_FAST (4.26 bpw, 13.55 GiB) GGUF release of Qwen 3.8 27B, custom-engineered for AMD Strix Halo (Ryzen AI Max+ 395 / Radeon 8060S) APUs.
🛠️ Deployment & Serving Code: github.com/julianmb/q38rocm
⚡ Includes: 1-click quickstart launcher (quickstart.sh), server launcher (run_server.sh), terminal TUI speedometer (chat_tui.py), NPU drafter orchestrator, and engine build scripts.
By combining ROCmFP4 block quantization, MTP (Multi-Token Prediction) Speculative Decoding, Asymmetric TurboQuant KV cache, and Mesa RADV Wave64 cooperative matrices (KHR_coopmat), this release achieves 30.56 – 36.04 tokens/second generation throughput on a single APU.
⚠️ Custom Engine Backend Required:
ROCmFP4is a custom ROCmFPX quantization layout designed for RDNA 3.5 / gfx1151 cooperative matrix hardware. It requires the ROCmFPX-enabledllama.cppengine fork (pinned build:e87d53e (213)). Upstream stockllama.cppor stock Ollama will fail to load ROCmFP4 GGUFs without this backend. See the q38rocm GitHub Repository for pre-compiled binaries and build instructions.
📊 Performance Benchmarks (AMD Ryzen AI Max+ 395)
Measured on AMD Ryzen AI Max+ 395 (40 CU Radeon 8060S @ 2.9 GHz, 128 GB LPDDR5X-8000 @ 273 GB/s, Linux 7.0, Mesa 26.0 RADV):
| Optimization Profile | Model Size | Unassisted Decode (Measured) | MTP Speculative Decode (Measured) | Speedup vs Baseline | TTFT (Prompt Eval) (Measured) |
|---|---|---|---|---|---|
Stock Q4_K_M (Baseline) |
15.92 GiB | 12.27 tok/s | N/A | 1.00× | 526.7 ms |
ROCmFP4_FAST (This Model) |
13.55 GiB | 14.02 tok/s | N/A | 1.14× | 468.3 ms |
ROCmFP4_FAST + Strict Greedy MTP |
13.55 GiB | 14.02 tok/s | 34.82 tok/s | 2.84× | 442.8 ms |
ROCmFP4_FAST + MTP (n6/p0.60) |
13.55 GiB | 14.02 tok/s | 30.56 – 34.82 tok/s | 2.50× – 2.84× | 439.4 ms |
ROCmFP4_FAST + Deep Spec (n7/p0.35) |
13.55 GiB | 14.02 tok/s | 🔥 36.04 tok/s | 🔥 2.94× | 445.8 ms |
💾 Context Scaling & Memory Footprint
Using Asymmetric TurboQuant KV cache (-ctk q8_0 -ctv turbo4):
| Context Window | Model Weights | TurboQuant KV Cache | Total RAM Footprint |
|---|---|---|---|
| 8K tokens | 13.55 GiB | 0.62 GiB | 14.17 GiB |
| 32K tokens | 13.55 GiB | 2.45 GiB | 16.00 GiB |
| 64K tokens | 13.55 GiB | 4.90 GiB | 18.45 GiB |
| 128K tokens | 13.55 GiB | 9.80 GiB | 23.35 GiB |
| 262K tokens (Full) | 13.55 GiB | 20.08 GiB | 33.63 GiB |
📥 Quick Download
# Using official HF CLI
hf download julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF Qwen3.8-27B-ROCmFP4-FAST.gguf --local-dir .
# Or using curl
curl -L "https://huggingface.co/julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF/resolve/main/Qwen3.8-27B-ROCmFP4-FAST.gguf" -o Qwen3.8-27B-ROCmFP4-FAST.gguf
SHA256 Checksum: fb89c78d2be91cdb68eaaaa45b1270710bf34aa721dc1f0b9e3aa7b98d2e1da9
🔍 GGUF Metadata Inspection
To inspect the quantization metadata, tensor architecture, and YaRN rope settings directly from the GGUF file:
from gguf import GGUFReader
reader = GGUFReader("Qwen3.8-27B-ROCmFP4-FAST.gguf")
print("Architecture:", reader.get_field("general.architecture"))
print("Quantization:", reader.get_field("general.file_type"))
print("Context Length:", reader.get_field("qwen3.context_length"))
print("Total Tensors:", len(reader.tensors))
🚀 How to Run
1. High-Throughput OpenAI API Server (MTP Speculation)
# Recommended environment for Strix Halo
export AMD_VULKAN_ICD=RADV
export VK_ICD_FILENAMES=/usr/share/vulkan/icd.d/radeon_icd.json
export HSA_OVERRIDE_GFX_VERSION=11.5.1
export GGML_HIP_ENABLE_UNIFIED_MEMORY=1
export RADV_PERFTEST="gpl,sam,nggc"
# Launch high-performance server
llama-server \
-m Qwen3.8-27B-ROCmFP4-FAST.gguf \
-dev Vulkan0 \
-ngl 99 \
-fa on \
-np 1 \
-ctxcp 0 \
-cram 16384 \
-c 32768 \
-b 2048 \
-ub 1024 \
-t 16 \
--poll 100 \
-ctk q8_0 \
-ctv turbo4 \
--port 8000 \
--spec-type draft-mtp \
--spec-draft-n-max 6 \
--spec-draft-p-min 0.60
2. Python (OpenAI SDK Client)
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="sk-no-key")
response = client.chat.completions.create(
model="qwen38-27b",
messages=[
{
"role": "system",
"content": "Reasoning effort is set to xhigh. Please think carefully through the task and prioritize correctness."
},
{
"role": "user",
"content": "Write a complete binary search tree implementation in Rust with insert and search methods."
}
],
temperature=0.7,
)
print(response.choices[0].message.content)
🔒 Limitations & Safety
- Custom Backend: Requires the ROCmFPX toolchain to execute.
- Hardware Target: Optimized specifically for AMD Strix Halo (RDNA 3.5 / gfx1151).
- Base Alignment: Inherits base safety characteristics and knowledge capabilities of Qwen 3.8 27B.
📜 License & Attribution
- Base Model: Qwen 3.8 27B by Alibaba Cloud
- Quantization & Optimizations: Apache 2.0 License.
- Community Research: NPU contention metrics referenced from ciru-ai's Strix Halo artifact.
- Downloads last month
- -
We're not able to determine the quantization variants.
Model tree for julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
Base model
Qwen/Qwen3.8-27BEvaluation results
- Peak Speculative Decode Speed on Strix Halo LLM Benchmark Suiteself-reported36.040
- Strict Lossless Greedy MTP Speed on Strix Halo LLM Benchmark Suiteself-reported34.820
- Base Unassisted Decode Speed on Strix Halo LLM Benchmark Suiteself-reported14.020
- Prompt Evaluation Latency (TTFT) on Strix Halo LLM Benchmark Suiteself-reported439.400
- Effective Bits Per Weight on Strix Halo LLM Benchmark Suiteself-reported4.260