-
incoai/GLM-5.3-Flash-DFlash2
Text Generation • 1B • Updated • 43.2k • 118 -
agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF
Text Generation • 177B • Updated • 79.9k • 84 -
julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF
Text Generation • 27B • Updated • 45.6k • 48 -
AesSedai/GLM-5.3-GGUF
753B • Updated • 1.33k • 7
Edward Paolo Guevarra PRO
PaoAI
AI & ML interests
None yet
Recent Activity
updated a model 1 day ago
PaoAI/Qwen3.8-27B-PaoAI-ROCmFP4-STRIX-BALANCED-GGUF posted an update 3 days ago
New on the Hub: **Qwen3.8-Flash-Next STRIX BALANCED-2.1** for AMD Strix Halo (Ryzen AI Max+ 395, 128 GB).
It's the same BALANCED-2 recipe with its always-on dense weights stored in 8-bit: 78.3 GB instead of 80.8, the same perplexity (−0.16 %, within error), and faster writing. It runs on a new ROCm/HIP engine built on @ilintar's Strix Halo llama.cpp branch plus one small fix of ours for the MTP check step.
Measured on one box with one fixed method: the full 262,144-token window checked at every depth (8K → 256K), a needle found at 260K, a 3-run coding exam graded by running the code (median 100/100), and a 55-task quality bench. Writing runs 17.5–46.7 t/s depending on depth and answer kind; reading runs 301–870 t/s.
👉 https://huggingface.co/PaoAI/Qwen3.8-Flash-Next-PaoAI-STRIX-BALANCED-2-GGUF
🛠 Engine: https://github.com/guevae2/paoai-qwen38fn-rocm-engine
Big thanks to @ilintar (Piotr Wilkin) for the strix-halo branch and ROCm runtime. It's his engine, and our part is one 3-line fix. Thanks also to @unsloth for the MTP draft sidecar and to Halogen for the 8-bit-dense idea. updated a model 3 days ago
PaoAI/Qwen3.8-Flash-Next-PaoAI-STRIX-BALANCED-2-GGUFOrganizations
None yet