Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
PaoAI 
posted an update 3 days ago
Post
79
New on the Hub: **Qwen3.8-Flash-Next STRIX BALANCED-2.1** for AMD Strix Halo (Ryzen AI Max+ 395, 128 GB).

It's the same BALANCED-2 recipe with its always-on dense weights stored in 8-bit: 78.3 GB instead of 80.8, the same perplexity (−0.16 %, within error), and faster writing. It runs on a new ROCm/HIP engine built on @ilintar 's Strix Halo llama.cpp branch plus one small fix of ours for the MTP check step.

Measured on one box with one fixed method: the full 262,144-token window checked at every depth (8K → 256K), a needle found at 260K, a 3-run coding exam graded by running the code (median 100/100), and a 55-task quality bench. Writing runs 17.5–46.7 t/s depending on depth and answer kind; reading runs 301–870 t/s.

👉 PaoAI/Qwen3.8-Flash-Next-PaoAI-STRIX-BALANCED-2-GGUF
🛠 Engine: https://github.com/guevae2/paoai-qwen38fn-rocm-engine

Big thanks to @ilintar (Piotr Wilkin) for the strix-halo branch and ROCm runtime. It's his engine, and our part is one 3-line fix. Thanks also to @unsloth for the MTP draft sidecar and to Halogen for the 8-bit-dense idea.
In this post