Download docs/PROTOCOL.md from EndlessChasing/Mamb2_8B_Recall: direct link, hf CLI and curl.
- Browser
- Download file 4.54 kB
-
https://huggingface.co/EndlessChasing/Mamb2_8B_Recall/resolve/main/docs/PROTOCOL.md
- Command line
-
hf download hf://EndlessChasing/Mamb2_8B_Recall/docs/PROTOCOL.md
-
curl -L -o PROTOCOL.md https://huggingface.co/EndlessChasing/Mamb2_8B_Recall/resolve/main/docs/PROTOCOL.md
Full precision Mamba2-8B Resurface control
Declared 2026-09-28 before fitting this repository's adapter. The purpose is to measure the effect of the same readout adapter and training budget on NVIDIA's uncompressed pure Mamba2-8B. The E8/W5 experiment is a historical comparator, not a training dependency. All metrics are reported at the actually measured precision and protocol, without assuming equality to native Megatron BF16.
Frozen model and intervention
Load nvidia/mamba2-8b-3t-4k revision
b915550c63ba9359f88f44d1f6a600d85af27302, checkpoint SHA256
47c2766f6aad89d73beafbeaecb334aab902d7370906d081764a90bb7a8bbbcb
and tokenizer SHA256
5862e2f71caf762bc9845662be5fec2867deb58d874568235a02a36c5111cd09.
Cast the 507 original BF16 tensors to FP16 in the native state-spaces Mamba2
runtime; freeze all 8,236,999,680 base parameters. Verify strict source tensor
mapping, model geometry and no shared embedding/output storage.
Use the same post-D adapter as the compressed experiment at all 56 native
gated RMSNorm inputs: y + sigmoid(w dot u+b) * g*(V@y) across 128 heads.
Its 224 tensors contain 1,154,104 FP32 trainable master parameters, exported
as FP16 for inference. Initialize V=0, g=1, w=0, b=-4; use soft sigmoid at
train and eval, no task switch, EMA, or new recurrent cache. The adapter is
external to the frozen base; check its identities and absence of base grads.
Training
Reuse the same exact data generator contract and seeds as the compressed
adapter: 1,536 TRAIN numeric bindings over three templates and N=16/64,
disjoint six-digit key/value intervals; full 256K CE on every answer suffix
token. Use seed 2026092803 and torch.randperm(1536) once for example order.
Each step pairs one MK example with a 512-token segment from the already
prepared 448 WikiText-2 TRAIN windows. Their historical manifest and file
SHA256 are pinned by the train command. Heldout and validation windows are
not included in training. Reusing this fixed TRAIN text gives the control the
same training exposures as the compressed arm.
Use a second, independently loaded uncompressed FP16 model as the frozen
prose teacher. This is the direct counterpart to the compressed experiment's
teacher, which was its own unadapted compressed base. The fixed per-step loss is
MK answer CE + 0.5 prose CE + 0.5 KL(own unadapted teacher || student) + 3 C,
where C is the same prose-only router closure, with budget 0.006 and excess
coefficient 10. Temperature 1, 511 prose targets per step, 64-token staged
head chunks, and one optimizer update after both task gradients.
AdamW FP32 masters: V/g LR 1e-4, w/b LR 3e-4, betas (.9,.999), eps 1e-8,
weight decay 0, clip norm 1. Multiplicative schedule
0.1+0.9*(1+cos(pi*j/1535))/2 at successful index j. Native FP16 forward,
block checkpointing, gradient scale 1024 with growth interval 2000. Overflow
retries the exact same pair without optimizer/master update, at most 8 retries
and 1,544 total attempts. Run 1,536 successful updates; the final step is
the only candidate. Small GPU smokes are discarded before formal training.
Evaluation and interpretation
After training, restore the actual serialized FP16 adapter onto the same uncompressed base. Evaluate baseline, adapter enabled and restored baseline in a single process, with identical tokenizer, MK prompts and full WikiText-2 validation windows. For MK use full 256K greedy generation up to 12 tokens, first standalone six-digit number, normal and target-removed controls, native prefill plus recurrent decode with fresh FP16 cache. For PPL use all 130 nonoverlapping reset windows and all 264,764 next-token targets. Record all raw scores, differences and file hashes; no early selection or task switch.
The earlier compressed arm's independent CONFIRM set has now been observed and uses the same three template families. Running this source control on it gives a matched historical comparison, not a newly untouched holdout. The WikiText-2 validation text also informed earlier development. We will not promote a source+adapter result to a general recall conclusion without new templates, longer distances and a truly untouched corpus.
Report all four observed arms explicitly: source FP16, source+adapter, compressed/readapted, compressed/readapted+adapter. The source and compressed teachers differ by construction. Report PPL and recall together; any cross-arm claim remains limited by shared test protocols and the historical numerical replay discrepancy in the compressed project.