quant-lab / README.md
kingjones777's picture
space: KingJones quant lab landing page
b2fe02d verified
|
Raw
History Blame Contribute Delete
936 Bytes
---
title: KingJones Quant Lab
emoji: πŸ”§
colorFrom: gray
colorTo: red
sdk: static
app_file: index.html
pinned: true
license: mit
short_description: ROCmFP4 & NVFP4 quants for Strix Halo β€” measured
tags:
- rocmfp4
- rocmfpx
- nvfp4
- strix-halo
- gfx1151
- ryzen-ai-max-395
- radeon-8060s
- amd
- rocm
- gguf
- quantization
- llama.cpp
- benchmark
---
# KingJones β€” ROCmFP4 & NVFP4 quant lab
Landing page for 18 published quantisations targeting AMD Strix Halo
(Ryzen AI Max+ 395, gfx1151, Radeon 8060S) and NVIDIA NVFP4.
Includes the measured finding that ROCmFP4 does **not** universally speed up decode,
the benchmark methodology behind every number, and the builds that were discarded.
- Models: <https://huggingface.co/kingjones777>
- Harness and raw results: <https://github.com/kingjones30/strix-halo-quant-lab>
- Quant format by [charlie12345/ROCmFPX](https://github.com/charlie12345/ROCmFPX)