New on the Hub: **Qwen3.8-Flash-Next STRIX BALANCED-2.1** for AMD Strix Halo (Ryzen AI Max+ 395, 128 GB).
It's the same BALANCED-2 recipe with its always-on dense weights stored in 8-bit: 78.3 GB instead of 80.8, the same perplexity (−0.16 %, within error), and faster writing. It runs on a new ROCm/HIP engine built on @ilintar's Strix Halo llama.cpp branch plus one small fix of ours for the MTP check step.
Measured on one box with one fixed method: the full 262,144-token window checked at every depth (8K → 256K), a needle found at 260K, a 3-run coding exam graded by running the code (median 100/100), and a 55-task quality bench. Writing runs 17.5–46.7 t/s depending on depth and answer kind; reading runs 301–870 t/s.
Big thanks to @ilintar (Piotr Wilkin) for the strix-halo branch and ROCm runtime. It's his engine, and our part is one 3-line fix. Thanks also to @unsloth for the MTP draft sidecar and to Halogen for the 8-bit-dense idea.