Spaces:
Running
Running
| title: KingJones Quant Lab | |
| emoji: π§ | |
| colorFrom: gray | |
| colorTo: red | |
| sdk: static | |
| app_file: index.html | |
| pinned: true | |
| license: mit | |
| short_description: ROCmFP4 & NVFP4 quants for Strix Halo β measured | |
| tags: | |
| - rocmfp4 | |
| - rocmfpx | |
| - nvfp4 | |
| - strix-halo | |
| - gfx1151 | |
| - ryzen-ai-max-395 | |
| - radeon-8060s | |
| - amd | |
| - rocm | |
| - gguf | |
| - quantization | |
| - llama.cpp | |
| - benchmark | |
| # KingJones β ROCmFP4 & NVFP4 quant lab | |
| Landing page for 18 published quantisations targeting AMD Strix Halo | |
| (Ryzen AI Max+ 395, gfx1151, Radeon 8060S) and NVIDIA NVFP4. | |
| Includes the measured finding that ROCmFP4 does **not** universally speed up decode, | |
| the benchmark methodology behind every number, and the builds that were discarded. | |
| - Models: <https://huggingface.co/kingjones777> | |
| - Harness and raw results: <https://github.com/kingjones30/strix-halo-quant-lab> | |
| - Quant format by [charlie12345/ROCmFPX](https://github.com/charlie12345/ROCmFPX) | |