Download docs/VERIFICATION.md from EndlessChasing/Mamb2_8B_Recall: direct link, hf CLI and curl.
- Browser
- Download file 5.41 kB
-
https://huggingface.co/EndlessChasing/Mamb2_8B_Recall/resolve/main/docs/VERIFICATION.md
- Command line
-
hf download hf://EndlessChasing/Mamb2_8B_Recall/docs/VERIFICATION.md
-
curl -L -o VERIFICATION.md https://huggingface.co/EndlessChasing/Mamb2_8B_Recall/resolve/main/docs/VERIFICATION.md
Independent verification of published PPL and MK
The published adapter and metrics at commit
e4137a43ef87cb95b70243ea2271f8e9b3d93d27 passed a fresh GPU replay on
2026-09-28. A new, clean clone was downloaded from public GitHub, all Python
source hashes were compared with that commit, and the numeric CONFIRM data
were regenerated. The adapter SHA-256 was
e8b2b4dfe69f8e85dc9e147c9aeaa3297cff14f4043e558bb1795ad476c1fca0.
The original reports were retained.
| Arm | WikiText-2 validation PPL | Normal MK | N=16 | N=64 | Target removed |
|---|---|---|---|---|---|
| Source FP16 | 7.334175947318572 | 147/384 | 113/192 | 34/192 | 0/384 |
| Source FP16 + Resurface | 7.052063515606311 | 365/384 | 191/192 | 174/192 | 0/384 |
The replay was exact for the measured outputs: both arms matched every one of their 130 per-window NLLs, with maximum absolute NLL error 0, and all 768 MK generated-token sequences, decoded strings, predictions and correctness flags. The final eight baseline probes after adapter removal also matched. Across the two arms, this covers 260 PPL windows and 1,536 MK generations. Run timestamps and GPU memory receipts are separate from this equality claim.
Normal MK increased by 218/384, or 56.77 percentage points: 222 previously wrong cases became correct and four previously correct cases became wrong. PPL decreased by 0.2821124317, or 3.8465%, on this validation protocol.
Checks independent of the inference loop
- A standard-library audit recomputed PPL as
exp(sum(NLL)/sum(targets))from every window and rescored each output string with the declared six-digit matcher. It checked exact input identities, paired gains/losses, summary arithmetic and the actual adapter file SHA. Its 17 self-tests exercise corrupt identities, wrong counts, numerical drift and receipt preservation. - A separate CPU pass used the pinned SentencePiece tokenizer to decode all 3,088 saved output records across the original/replay reports, including restoration probes. It regenerated and checked their prompt token hashes and expected answers. All 520 PPL window receipts across both reports and arms were checked against newly loaded, pinned WikiText tokens. There were zero mismatches.
- Regenerating the 1,536 numeric TRAIN examples and 768 CONFIRM rows found zero shared case IDs, prompts, record keys or record values between those splits. The training receipt contains exactly one successful update for each TRAIN example and six overflow retries.
- The pinned NVIDIA model revision and WikiText revision resolved on the Hugging Face Hub. Official model/tokenizer LFS SHA-256 values matched the runtime pins. The GPU loader verifies the 16,474,189,490-byte checkpoint before strict tensor mapping into the native model.
The actual environment was Python 3.10.12, PyTorch 2.11.0+cu128, Mamba-SSM 2.3.2.post1, Triton 3.6.0, NumPy 1.26.4, datasets 4.8.5 and SentencePiece 0.2.1 on an RTX PRO 6000 Blackwell Server Edition, with TF32 matmul disabled.
Receipts and reproduction
- Fresh GPU replay
- Independent metric and exact-output audit
- Tokenizer and regenerated-input checks
- Clean-clone origin, source hashes, command and environment
- Pinned official Hub metadata
Recheck the saved reports without a GPU or PyTorch:
python3 scripts/verify_replay.py \
--published reports/source_resurface_v1_confirm_full.json \
--replay reports/verification_replay_v1.json \
--adapter artifacts/source_resurface_v1/adapter_fp16.pt \
--output reports/my_verification_audit.json
An audit exit code of zero means internally valid reports; inspect
exact_replay and status for equality. Valid numerical differences are
reported explicitly. The script refuses to overwrite an existing receipt.
For a fresh GPU run, follow REPRODUCE.md with a new output path.
The review found a documentation portability defect: numeric manifests embed the Python build string, so their whole-file SHA can differ even for identical data. The reproduction commands now use each newly generated manifest's SHA; the loaders still validate the frozen generator, tokenizer, protocol and all raw/token content. The docs also clarify the separate native Mamba extension installation. Neither change alters the measured model or scoring protocol.
Scope
PPL uses WikiText-2 validation, 130 reset windows of up to 2,048 next-token targets (264,764 targets total), with parallel SSD scan precision. The model is the official BF16 source cast to FP16 in the native runtime; this does not establish native Megatron BF16 parity or tokenwise FP16-state rounding.
MK uses three synthetic numeric templates with 16 or 64 bindings and prompt lengths of 257-1,236 tokens. The CONFIRM instances and validation corpus had been observed during earlier development. Exact replay confirms these reported scores, not unseen-template or full-4K recall generalization. Target-removed 0/384 means the absent answer was not recovered; it is not a refusal-calibration metric. Base freezing was checked by identities, version counters and gradients, not a complete post-training bytewise weight scan.