jcbtc commited on
Commit
97082b0
·
verified ·
1 Parent(s): 2528934

Pin ROCmFPX runner and add build instructions

Browse files
Files changed (1) hide show
  1. README.md +49 -3
README.md CHANGED
@@ -41,7 +41,7 @@ The goal is simple: keep Step 3.7 Flash useful at 256K context, keep the quality
41
 
42
  Use this if you want the Step 3.7 behavior profile, MTP support, and a much smaller local footprint than the stock GGUF Q3_K_L or ROCmFP4 STRIX_LEAN builds.
43
 
44
- > Runtime note: these GGUFs use ROCmFPX tensor types. They are intended for the Chadrock / ROCmFPX llama.cpp runner family, not stock upstream llama.cpp.
45
 
46
  ## Why This One
47
 
@@ -161,6 +161,49 @@ Download it from the Files tab or directly from:
161
  https://huggingface.co/jcbtc/Step-3.7-Flash-ROCmFPX-Q3-QualityPlus/resolve/main/step37-native-tool-response-template.jinja
162
  ```
163
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
164
  ## Recommended Serving Profile
165
 
166
  The locally tested long-context profile:
@@ -186,7 +229,7 @@ chat template: Step native tool_response template with protocol-boundary escapin
186
  Example shape:
187
 
188
  ```bash
189
- /path/to/ROCmFPX/build/bin/llama-server \
190
  -m Step-3.7-Flash-ROCmFPX-Q3-QualityPlus-00001-of-00009.gguf \
191
  --alias step-3.7-flash-rocmfpx-q3-qualityplus \
192
  --host 127.0.0.1 \
@@ -235,6 +278,8 @@ That matters for real agents because Step 3.7 can otherwise confuse tool output
235
 
236
  ## Build Notes
237
 
 
 
238
  The QualityPlus policy used here:
239
 
240
  - huge `ffn_*_exps` tensors: `q3_0_rocmfpx`
@@ -250,10 +295,11 @@ Converter-reported size: `83726.08 MiB / 3.57 BPW`, 9 shards.
250
  - Base model: [`stepfun-ai/Step-3.7-Flash`](https://huggingface.co/stepfun-ai/Step-3.7-Flash)
251
  - MTP draft GGUF source: [`notSnix/Step-3.7-Flash-MTP-Draft-GGUF`](https://huggingface.co/notSnix/Step-3.7-Flash-MTP-Draft-GGUF)
252
  - ROCmFPX creator: Charlie, `charlie12345` / `@italianclownz`, [`charlie12345/ROCmFPX`](https://github.com/charlie12345/ROCmFPX)
 
253
  - Quantization, the ROCmFPX Step 3.7 Q3 QualityPlus recipe, Strix Halo profile, and local benchmark work: Crown / Ciru
254
 
255
  ## Caveats
256
 
257
- - This is a custom ROCmFPX GGUF release. Use a compatible ROCmFPX/Chadrock llama.cpp runner.
258
  - Quality numbers are local Strix Halo measurements and depend on runtime, chat template, KV type, and MTP settings.
259
  - The model is strong but not perfect at autonomous email/message side effects; it can be cautious and ask for subject/body/recipient details instead of sending with inferred defaults.
 
41
 
42
  Use this if you want the Step 3.7 behavior profile, MTP support, and a much smaller local footprint than the stock GGUF Q3_K_L or ROCmFP4 STRIX_LEAN builds.
43
 
44
+ > **Required runtime:** these GGUFs do **not** run on stock upstream llama.cpp. They use ROCmFPX tensor types such as `q3_0_rocmfpx` plus Chadrock/ROCmFPX serving support for Step MTP. Build the pinned Ciru ROCmFPX runner below before trying to load the model.
45
 
46
  ## Why This One
47
 
 
161
  https://huggingface.co/jcbtc/Step-3.7-Flash-ROCmFPX-Q3-QualityPlus/resolve/main/step37-native-tool-response-template.jinja
162
  ```
163
 
164
+ ## Required ROCmFPX Runner
165
+
166
+ This model is tied to the Charlie/Ciru ROCmFPX llama.cpp runner family. A stock `llama-server` will not understand the ROCmFPX tensor types in these shards and will not reproduce the MTP serving behavior used for the benchmark rows.
167
+
168
+ Use the pinned Ciru runner:
169
+
170
+ ```text
171
+ repo: https://github.com/ciru-ai/ROCmFPX
172
+ current recommended pin: 221402af8574faf652b101b6afe225a3f329561f
173
+ branch at time of pin: main
174
+ upstream lineage: charlie12345/ROCmFPX
175
+ ```
176
+
177
+ The earlier Chadrock v2 speed-runner tag remains useful for historical comparison:
178
+
179
+ ```text
180
+ tag: chadrockv2-runner-20260622
181
+ commit: 7aa484a2f0a504dc612a3d74a068024f3e6d6353
182
+ ```
183
+
184
+ The Q3 QualityPlus Step 3.7 rows on this card were validated with the Chadrock/ROCmFPX runner path on AMD Ryzen AI Max+ 395 / Strix Halo. For fresh installs, use the current Ciru pin above unless you are reproducing an older benchmark exactly.
185
+
186
+ Build the runner:
187
+
188
+ ```bash
189
+ git clone https://github.com/ciru-ai/ROCmFPX.git
190
+ cd ROCmFPX
191
+ git checkout 221402af8574faf652b101b6afe225a3f329561f
192
+
193
+ env JOBS="$(nproc)" \
194
+ CMAKE_HIP_ARCHITECTURES=gfx1151 \
195
+ ROCMFPX_DECODE_TUNE=stable \
196
+ scripts/build-strix-rocmfp4-mtp.sh llama-server llama-bench
197
+ ```
198
+
199
+ The server binary should be:
200
+
201
+ ```text
202
+ ./build-strix-rocmfp4/bin/llama-server
203
+ ```
204
+
205
+ If the model load fails with an unknown GGUF tensor type, you are using the wrong runner.
206
+
207
  ## Recommended Serving Profile
208
 
209
  The locally tested long-context profile:
 
229
  Example shape:
230
 
231
  ```bash
232
+ ./build-strix-rocmfp4/bin/llama-server \
233
  -m Step-3.7-Flash-ROCmFPX-Q3-QualityPlus-00001-of-00009.gguf \
234
  --alias step-3.7-flash-rocmfpx-q3-qualityplus \
235
  --host 127.0.0.1 \
 
278
 
279
  ## Build Notes
280
 
281
+ These are model-build notes, not runner-build instructions. Build the pinned ROCmFPX runner in the section above before serving the GGUFs.
282
+
283
  The QualityPlus policy used here:
284
 
285
  - huge `ffn_*_exps` tensors: `q3_0_rocmfpx`
 
295
  - Base model: [`stepfun-ai/Step-3.7-Flash`](https://huggingface.co/stepfun-ai/Step-3.7-Flash)
296
  - MTP draft GGUF source: [`notSnix/Step-3.7-Flash-MTP-Draft-GGUF`](https://huggingface.co/notSnix/Step-3.7-Flash-MTP-Draft-GGUF)
297
  - ROCmFPX creator: Charlie, `charlie12345` / `@italianclownz`, [`charlie12345/ROCmFPX`](https://github.com/charlie12345/ROCmFPX)
298
+ - Pinned public runner fork and build recipe: [`ciru-ai/ROCmFPX`](https://github.com/ciru-ai/ROCmFPX), current recommended pin `221402af8574faf652b101b6afe225a3f329561f`
299
  - Quantization, the ROCmFPX Step 3.7 Q3 QualityPlus recipe, Strix Halo profile, and local benchmark work: Crown / Ciru
300
 
301
  ## Caveats
302
 
303
+ - This is a custom ROCmFPX GGUF release. It requires the compatible ROCmFPX/Chadrock llama.cpp runner; stock llama.cpp is not expected to load it.
304
  - Quality numbers are local Strix Halo measurements and depend on runtime, chat template, KV type, and MTP settings.
305
  - The model is strong but not perfect at autonomous email/message side effects; it can be cautious and ask for subject/body/recipient details instead of sending with inferred defaults.