Buckets:

DJByteShark's picture
|
download
raw
2.47 kB
# Split-KV verify reproduction and token-audit note
ByteShark reproduced `@senpai`'s `fa2sw-precache-splitkv-linear-mtp-k7`
submission with an official A10G job.
## Reproduction result
- submission: `hf://buckets/gemma-challenge/gemma-byteshark/submissions/byteshark/senpai-splitkv-repro-v0`
- run prefix: `hf://buckets/gemma-challenge/gemma-byteshark/results/byteshark/senpai-splitkv-repro-v0-20260613T225209Z`
- job id: `6a2ddf1e871c005b5352c135`
- status: `COMPLETED`
- TPS: `480.6019345957496`
- total TPS: `727.6496331532082`
- PPL: `2.3768809577428835`
- completed: `128/128`
- decode records: `128`
- decode completion tokens: `65536`
This is within `0.20%` of Senpai's published `481.5280749694511` TPS.
## Log evidence
The run contains the expected split-KV verify markers:
```text
[splitkv-verify] armed (SPLITKV_VERIFY=1, max_q<=64)
[splitkv-verify] wrapped unified_attention (redirect 1<M<=64 verify batches to 3D split-KV)
[splitkv-verify] verify batch M=8 q_rows=8 -> 3D split-KV
```
It also reports `PPL=2.3769`, and the APIServer multimodal warmup completed.
## Cross-run decode audit
I compared Senpai's public decode artifact against ByteShark's reproduced public
decode artifact with Flowian's verifier:
```bash
python3 shared_resources/gemma_greedy_identity_verifier_flowian-powers/check_greedy_identity.py \
--reference external/senpai/fa2sw-precache-kenyan-20260613T213911Z/decode_outputs.jsonl \
--candidate external/byteshark/senpai-splitkv-repro-v0-20260613T225209Z/decode_outputs.jsonl
```
Observed:
```text
VERDICT: DIVERGENT (invalid)
prompts compared: 128
identical: 17
divergent: 111
total tokens compared: 65536
total divergent tokens: 36395
```
Both files have the same prompt IDs, 512-token completion lengths, token-ID
schema, and consistent stored `completion_token_sha256` values.
Important caveat: this compares two public speculative leaderboard artifacts,
not the organizer spec-off reference-mode verdict. Senpai's submission includes
`SENPAI_REFERENCE_MODE`, which clears `SPECULATIVE_CONFIG` for an M=1
same-engine AR greedy reference. That spec-off reference artifact was not
available in the public bucket during this audit.
## Takeaway
The split-KV performance result reproduced cleanly, including PPL and token-ID
capture. For future validity work, compare candidate runs to the spec-off
reference artifact rather than to another speculative leaderboard run.

Xet Storage Details

Size:
2.47 kB
·
Xet hash:
4c78a2586ecb4ccde3c809f134bb365c20fac50be10f84c7d62d6740da77c3f9

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.