# Runtime validation This release was gated on successful loading, server startup, and real token generation. It was not gated on coding-answer quality. ## Structural checks Run from the release workspace root: ```powershell $Python = '.\.venv\Scripts\python.exe' $Validator = '.\tools\validate_rivetcoder_gguf.py' $Llama = '.\work\llama.cpp-rivetcoder' $Source = '.\RivetCoder-9B-A4B' $Package = '.\RivetCoder-9B-A4B-GGUF' & $Python $Validator "$Package\RivetCoder-9B-A4B-Q8_0.gguf" ` --llama-cpp-dir $Llama ` --hf-model $Source ` --expected-file-type Q8_0 ` --strict-extra-tensors ` --report "$Package\provenance\q8-structure.json" & $Python $Validator "$Package\RivetCoder-9B-A4B-Q4_K_M.gguf" ` --llama-cpp-dir $Llama ` --hf-model $Source ` --expected-file-type Q4_K_M ` --strict-extra-tensors ` --report "$Package\provenance\q4km-structure.json" ``` Both reports must contain `"ok": true`, 506 tensors, no missing or extra tensors, no warnings, and zero routing-control mismatches. ## CLI generation check After extracting `bundle/rivetcoder-llama-windows-cuda-sm120.zip`: ```powershell .\llama-cli.exe ` --model .\RivetCoder-9B-A4B-Q4_K_M.gguf ` --n-gpu-layers all ` --ctx-size 512 ` --batch-size 64 ` --ubatch-size 64 ` --predict 16 ` --prompt 'Reply with OK only.' ` --temperature 0 ` --single-turn ` --simple-io ` --no-warmup ` --perf ``` Success requires model load, at least one generated token, and process exit code `0`. Repeat with the Q8_0 file. ## Server check ```powershell $server = Start-Process ` -FilePath '.\llama-server.exe' ` -ArgumentList @( '--model', '.\RivetCoder-9B-A4B-Q4_K_M.gguf', '--alias', 'RivetCoder-9B-A4B', '--host', '127.0.0.1', '--port', '8080', '--n-gpu-layers', 'all', '--ctx-size', '4096', '--parallel', '1', '--batch-size', '512', '--ubatch-size', '256', '--flash-attn', 'on', '--reasoning-budget', '0' ) ` -PassThru ` -WindowStyle Hidden Invoke-RestMethod http://127.0.0.1:8080/health $body = @{ model = 'RivetCoder-9B-A4B' messages = @(@{ role = 'user'; content = 'Say OK.' }) temperature = 0 max_tokens = 16 } | ConvertTo-Json -Depth 8 Invoke-RestMethod ` -Method Post ` -Uri http://127.0.0.1:8080/v1/chat/completions ` -ContentType application/json ` -Body $body Stop-Process -Id $server.Id ``` Success requires `GET /health` to return `{"status":"ok"}`, a non-error completion response containing generated tokens, and no CUDA/graph/tensor errors in the server log. ## Recorded results | Artifact | Load/generate | Structure | SHA-256 | |---|---|---|---| | BF16 conversion | pass, exit 0 | pass | `f51dfc1cb82b281156a09360bdbe8fe95734334e01d0d6b6c2df4a27b2619aea` | | Q8_0 | pass, exit 0 | pass | `6e71dd349df93f734b89fd70f2d0590037ddd3bdc11a4e9533991a0eab47b85e` | | Q4_K_M | pass, exit 0 | pass | `980adc81f64b52bc71228abd0e2bed02a1b6a00bdda9afbfd6d5e0c124ac5600` | The BF16 intermediate is intentionally not part of the public quantized release.