memra-bench / CONFIGS.md
Avifenesh's picture
27B ST serve config: + BW24_MMQ_F8F4=1 (f8f4 prefill tile adoption)
ec242fd verified
|
Raw
History Blame Contribute Delete
5.78 kB

bw24-bench — exact configs for v0.8.0 speculative-decoding board

Reproducibility configs for the speculative-decoding (ST) rows added in v0.8.0. These correspond to the README Performance section's safetensors-path models (Qwen3.5-9B ST and Qwen3.6-27B ST).


Qwen3.5-9B NVFP4 ST (modelopt safetensors)

Model: /data/ai-ml/hf-models/qwen35-9b-nvfp4-st-modelopt/ (modelopt NVFP4 safetensors checkpoint)

Plain decode (no speculation)

# No flags — tuned path is the default
./target/release/run-gen /data/ai-ml/hf-models/qwen35-9b-nvfp4-st-modelopt/ --prompt "<prompt>"

Speculative decode (MTP + FR-Spec trim)

Short-code (p1):

BW24_SPEC_K=2 \
BW24_SPEC_PMIN=0.3 \
BW24_FRSPEC_TRIM=trims/frspec-9bst-modelopt-32768.gguf \
./target/release/run-spec /data/ai-ml/hf-models/qwen35-9b-nvfp4-st-modelopt/ \
  --prompt "$(cat prompts/p1-code-short.txt)"

Medium-code (p2) and Agentic-long (p3):

BW24_SPEC_K=3 \
BW24_SPEC_PMIN=0.3 \
BW24_FRSPEC_TRIM=trims/frspec-9bst-modelopt-32768.gguf \
./target/release/run-spec /data/ai-ml/hf-models/qwen35-9b-nvfp4-st-modelopt/ \
  --prompt "$(cat prompts/p2-code-medium.txt)"  # or p3-agentic-long.txt

Qwen3.6-27B NVFP4 ST (NVIDIA official safetensors)

Model: /data/ai-ml/hf-models/nvidia-qwen36-27b-nvfp4/ (NVIDIA's official release)

Plain decode (no speculation)

BW24_NV_W4=1 BW24_MMQ_F8F4=1 \
./target/release/run-gen /data/ai-ml/hf-models/nvidia-qwen36-27b-nvfp4/ --prompt "<prompt>"

Speculative decode (MTP + FR-Spec trim + HPOST)

Short-code (p1):

BW24_NV_W4=1 BW24_MMQ_F8F4=1 \
BW24_SPEC_K=3 \
BW24_SPEC_HPOST=1 \
BW24_SPEC_PMIN=0.4 \
BW24_FRSPEC_TRIM=trims/frspec-corpus-32768.gguf \
./target/release/run-spec /data/ai-ml/hf-models/nvidia-qwen36-27b-nvfp4/ \
  --prompt "$(cat prompts/p1-code-short.txt)"

Medium-code (p2):

BW24_NV_W4=1 BW24_MMQ_F8F4=1 \
BW24_SPEC_K=3 \
BW24_SPEC_HPOST=1 \
BW24_SPEC_PMIN=0.3 \
BW24_FRSPEC_TRIM=trims/frspec-corpus-32768.gguf \
./target/release/run-spec /data/ai-ml/hf-models/nvidia-qwen36-27b-nvfp4/ \
  --prompt "$(cat prompts/p2-code-medium.txt)"

Agentic-long (p3):

BW24_NV_W4=1 BW24_MMQ_F8F4=1 \
BW24_SPEC_K=3 \
BW24_SPEC_HPOST=1 \
BW24_SPEC_PMIN=0.4 \
BW24_FRSPEC_TRIM=trims/frspec-corpus-32768.gguf \
./target/release/run-spec /data/ai-ml/hf-models/nvidia-qwen36-27b-nvfp4/ \
  --prompt "$(cat prompts/p3-agentic-long.txt)"

IMPORTANT NOTE: The p3-agentic-long column carries a degeneration caveat (see rig5090.jsonl tag p3-degeneration-audit): greedy continuations on this prompt partially degenerate into repetition on both engines symmetrically. The ratios between engines remain valid, but absolute tok/s values overstate real agentic throughput. A replacement prompt is pending.


llama.cpp comparison configs

These are the exact commands used to benchmark llama.cpp on the same hardware (RTX 5090 Laptop). Both engines built and measured in the same session.

Build info

  • Commit: b9743 (c57607016)
  • CUDA: 13.1
  • Compile flags: GGML_CUDA_FA=ON GGML_CUDA_GRAPHS=ON GGML_CUDA_FA_ALL_QUANTS=ON, FORCE_MMQ=OFF
  • Architecture: sm_120a SASS
  • Binary path: /home/avifenesh/projects/llama.cpp/build/bin/
  • Required env: export GGML_CUDA_GRAPH_OPT=1 before every run

llama-bench (raw pp/tg, no MTP)

9B NVFP4:

export GGML_CUDA_GRAPH_OPT=1
/home/avifenesh/projects/llama.cpp/build/bin/llama-bench \
  -m /data/ai-ml/hf-models/qwen35-9b-nvfp4-gguf/Qwen3.5-9B-NVFP4-MTP-GGUF.gguf \
  -ngl 999 -fa 1 -ctk q8_0 -ctv q5_1 -p 2048 -n 128 -d 0,4096,8192 -r 5
# MEASURED: pp512=6220 t/s, tg128=126.6 t/s (raw decode, no MTP)

27B NVFP4:

export GGML_CUDA_GRAPH_OPT=1
/home/avifenesh/projects/llama.cpp/build/bin/llama-bench \
  -m /data/ai-ml/hf-models/qwen36-27b-nvfp4-mtp/Qwen3.6-27B-NVFP4-Q4_K_M-mtp.gguf \
  -ngl 999 -fa 1 -ctk q8_0 -ctv q5_1 -p 2048 -n 128 -d 0,4096,8192 -r 5
# MEASURED: pp512=1980 t/s, tg128=42.1 t/s (raw decode, no MTP)

llama-server (with MTP speculative decoding)

27B NVFP4 with MTP (peak config):

export GGML_CUDA_GRAPH_OPT=1
/home/avifenesh/projects/llama.cpp/build/bin/llama-server \
  -m /data/ai-ml/hf-models/qwen36-27b-nvfp4-mtp/Qwen3.6-27B-NVFP4-Q4_K_M-mtp.gguf \
  -md /data/ai-ml/hf-models/qwen36-27b-nvfp4-mtp/mtp-Qwen3.6-27B-NVFP4.gguf \
  -ngl 999 -fa on -ctk q8_0 -ctv q5_1 -c 65536 --parallel 1 \
  --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.2 \
  --jinja --temp 0.6 --top-p 0.95 --top-k 20 --host 127.0.0.1 --port 8099
# MEASURED: 66.6 t/s decode with MTP (accept rate 0.762, mean accept len 3.29)

9B NVFP4 serve (no separate MTP draft on disk):

export GGML_CUDA_GRAPH_OPT=1
/home/avifenesh/projects/llama.cpp/build/bin/llama-server \
  -m /data/ai-ml/hf-models/qwen35-9b-nvfp4-gguf/Qwen3.5-9B-NVFP4-MTP-GGUF.gguf \
  -ngl 999 -fa on -ctk q8_0 -ctv q5_1 -c 65536 --parallel 1 \
  --jinja --temp 0.6 --top-p 0.95 --top-k 20 --host 127.0.0.1 --port 8099

Protocol notes

  • N=3 runs minimum, report the median
  • 256 generated tokens, temperature 0 (greedy)
  • Same session: both engines measured within the same thermal regime, no clock-speed drift
  • Speculative exactness gate: K=1..8 self-consistency check pins output token-identical to plain greedy decode
  • Numeric config: all v0.8.0 numbers use BW24_FA_V2=1 (tile-batched online-softmax FA decode, the engine default since 2026-07-08)

Full protocol details in the main repo: research/benchmarks.md and docs/COMPETITOR-SETUP.md.