Mamb2_8B_Recall_SQ3.25
Official WikiText-2 test PPL — Resurface on: 7.77130287.
| Published configuration | Resurface adapter | Official WT2 test PPL |
|---|---|---|
| Same frozen weights and SQ3.25 state | Off | 8.11328914 |
| Same frozen weights and SQ3.25 state | On | 7.77130287 |
147 reset windows · 300,963 next-token targets, including the final partial
window. Measured with the published StatePPLQuantV10 32/32/64 codec: carry is
requantized after every token, with FP16 computation and FP32 logits.
The checkpoint was fixed before test evaluation; no test-driven training or selection.
Packed 3.25-bit recurrent state plus a freshly trained Resurface adapter
for the pure nvidia/mamba2-8b-3t-4k
language model. Full official WT2 test PPL is 7.77130287, versus 8.11328914
without the adapter, using 27.1797 MiB of persistent state and table storage
per sequence at batch one. Historical normal MK is 271/384 (70.57%); it was
not rerun for this test update.
SQ3.25 describes recurrent state. The 8B model weights remain FP16. This
repository distributes the adapter, calibration table, custom runtime and
verification evidence. Download the original NVIDIA base separately. It is
not a standalone 8B checkpoint, a 3.25-bit weight model, or a Transformers/PEFT
from_pretrained package.
This is the SQ3.25 member of the Mamb2_8B_Recall family. Its adapter was trained fresh on the fixed quantized-state base; it is distinct from the original Mamb2_8B_Recall adapter.
Full official WikiText-2 test PPL — fixed published checkpoint
| Configuration | WikiText-2 test PPL |
|---|---|
| Published SQ3.25 state, without Resurface | 8.11328914 |
| Same frozen FP16 base and state + released V11 Resurface | 7.77130287 |
Both complete arms score 300,963 next-token targets in 147 reset windows,
including the last window of 1,955 targets. The adapter lowers pooled-token PPL
by 4.2151%. The weights, 32/32/64 state table, final adapter and measured
runtime are unchanged from v0.1.0-q325-resurface. No training, recalibration,
checkpoint/candidate selection, or test-dependent stopping was performed.
Protocol: pinned Salesforce/wikitext, wikitext-2-raw-v1, official test
split at revision b08601e04326c79dfdd32d625aee71d232d685c3; documents joined
by two newlines and NVIDIA SentencePiece without automatic BOS/EOS. Each window
scores at most 2,048 next-token targets with one boundary-token overlap; state
is requantized every token. PPL is exp(total NLL / 300963) from FP16 compute
and FP32 logits.
Data scope: adapter training used separate official TRAIN data. The test was scored after publication with the fixed released checkpoint and without test-driven selection. Earlier project experiments, including 2.7B and E8/W5, used the official WT2 test text, so this is not an untouched test for the whole project. Source-model pretraining contamination and cross-split duplication have not been audited. No S16 test control or new MK evaluation was run here.
The exact comparison, adapter-free arm, adapted arm, and CPU audit contain all per-window scores/token hashes and the recorded tensor/storage bindings, without raw token IDs. The CPU audit verifies token coverage, pooled NLL/PPL arithmetic, the published runtime and 507-weight/224-adapter bindings, and adapter-removal reset/cache restoration. It does not recompute GPU logits. See the reproduction guide, frozen test protocol, and evidence inventory.
This is a documentation/evaluation overlay. The original release tag, adapter,
state table, runtime source and sealed release/ files remain unchanged.
Historical validation and MK CONFIRM results remain separate below; they are
not WT2-test MK or newly collected test recall measurements.
Historical validation and published MK CONFIRM
| Configuration | WikiText-2 validation PPL | Normal MK | N=16 | N=64 |
|---|---|---|---|---|
| Original S16, archived reference | 7.334322057221965 | 146/384 (38.02%) | — | — |
| Fixed Q3.25 state, no adapter | 8.186186562837207 | 32/384 (8.33%) | 30/192 | 2/192 |
| Q3.25 + released Resurface | 7.855569605864836 | 271/384 (70.57%) | 181/192 | 90/192 |
| Q3.25 after adapter removal | 8.186186562837207 | 32/384 (8.33%) | 30/192 | 2/192 |
PPL covers 130 reset windows of up to 2,048 predicted tokens, with 264,764 targets in the pinned WikiText-2 validation split. MK uses 384 normal prompts and 384 paired target-removed diagnostic prompts. Both Q3.25 configurations match the removed target 0/384 times; this is not an abstention score. The normal-MK denominator excludes those controls.
Against the fixed Q3.25 base, the adapter improves PPL 4.0387% and normal MK 62.2396 percentage points: 242 paired gains and three regressions. The paired bootstrap 95% interval for the MK change is [+57.2917, +67.1875] percentage points. PPL remains 7.11% higher than the archived S16 reference, which was not rerun in this experiment.
All nine predefined quality and integrity gates passed, including the PPL threshold, positive paired recall improvement, exact control replay and unchanged persistent cache. Removing the adapter reproduced all 130 baseline window NLLs and all 768 generated sequences exactly, including reset/cache receipts. An independent CPU audit of the recorded full evidence also passed; that audit does not rerun GPU logits. See full results, raw comparison and full audit.
State format and memory
Each 128-coordinate recurrent-state row stores 32 INT8 coordinates, 32 INT4 coordinates and 64 zero-carry coordinates, plus two FP16 scales. That is 52 bytes per row, or 3.25 bits per coordinate including scales. The custom runtime uses the packed carry for every token, including prompt processing. The selected coordinate table is fixed before adapter training.
| Persistent component, batch one | Bytes | Size |
|---|---|---|
| SSM cache | 23,855,104 | 22.7500 MiB |
| Convolution cache | 4,587,520 | 4.3750 MiB |
| One shared static table | 57,344 | 0.0547 MiB |
| Cache plus table | 28,499,968 | 27.1797 MiB |
| Separate FP16 adapter payload | 2,308,208 | 2.3082 MB |
| Cache, table and adapter | 30,808,176 | 29.3810 MiB |
| Original FP16 model weights, separate | 16,473,999,360 | 16.4740 GB |
The cache/table figure is 76.64% smaller than the archived S16 cache of 116.3750 MiB. Static table and adapter storage can be shared across requests; each is counted once here. These are tensor payloads, excluding temporary computation, allocator reserve and training teacher/optimizer state. They are not total GPU memory requirements. The serialized adapter file is 2,375,743 bytes; its 1,154,104 FP16 parameters occupy 2,308,208 bytes when loaded.
Download and verify
With the Hugging Face CLI installed:
hf download EndlessChasing/Mamb2_8B_Recall_SQ3.25 \
--local-dir Mamb2_8B_Recall_SQ3.25
cd Mamb2_8B_Recall_SQ3.25
python3 scripts/release_state_resurface_v11.py --verify release
The standard-library verifier checks the exact adapter, table, raw evidence,
internal checksums and unchanged measured runtime files. Verification does
not require CUDA or download the base. The original 18-file release bundle
is preserved under release/; runtime code, scripts and notices are supplied
alongside it. This verification checks published evidence rather than running
a new quality evaluation.
Download the separate source checkpoint and tokenizer at the pinned revision:
hf download nvidia/mamba2-8b-3t-4k \
release/mp_rank_00/model_optim_rng.pt \
mt_nlg_plus_multilingual_ja_zh_the_stack_frac_015_256k.model \
--revision b915550c63ba9359f88f44d1f6a600d85af27302 \
--local-dir models/source
The loader verifies the checkpoint and tokenizer SHA-256 values, then casts the official BF16 weights to FP16. Native NVIDIA Megatron BF16 equivalence is not established by these measurements.
Environment and inference
The measured environment was Linux, Python 3.10.12, PyTorch 2.11.0+cu128, Triton 3.6.0, Mamba-SSM 2.3.2.post1, NumPy 1.26.4, datasets 4.8.5 and SentencePiece 0.2.1, on an RTX PRO 6000 Blackwell Server Edition. Install the matching CUDA/PyTorch native Mamba and causal-convolution dependencies first, then install this package from the downloaded root:
python -m pip install --no-deps -e .
python scripts/infer_state_resurface_v11.py \
--bundle release \
--source-dir models/source \
--prompt 'The capital of France is' \
--max-new-tokens 32
The entrypoint checks recorded backend versions, external source hashes,
precision flags and the fixed 16-warp RMSNorm configuration. An incompatible
backend fails before model loading. pyproject.toml contains broad dependency
ranges, not an exact environment lock; other GPU architectures are not validated.
See the release guide.
This wrapper provides greedy batch-one generation with prompts up to 4,096
tokens. Its separate GPU smoke reproduced one archived 263-token MK prompt
and all 12 continuation token IDs, with the same cache bytes. That check
validates the wrapper on that case; it does not repeat or extend the full
PPL/MK evaluation. The example prompt above is a usage example, not a reported
quality measurement. Native InferenceParams, variable-length batching and
a serving integration are not provided.
To check the package through the inference entrypoint without torch or CUDA:
python3 scripts/infer_state_resurface_v11.py --bundle release --verify-only
Training and limitations
The original 8,236,999,680 base parameters remain frozen. After fixing the selected state table, a fresh post-D Resurface-style adapter completed 1,536 successful TRAIN updates, with five numerical overflow retries. Only the final TRAIN-derived FP16 export was evaluated; validation/CONFIRM did not select a checkpoint. Training/export forward parity was checked at 128 and 512 tokens, with separately audited gradient fixtures.
The validation corpus and CONFIRM prompt families have historical exposure. The official test text also appeared in earlier project experiments, as scoped in the fixed-checkpoint test section above. The results establish performance on this specified protocol. They do not establish unseen-template recall, full-4K recall quality, other-language performance or broad language-model generalization. The wrapper's 4K prompt limit is an interface constraint, not a demonstrated 4K recall result. The historical validation source identity/version/gradient guards did not constitute a complete post-run bytewise base audit. The new fixed-checkpoint test records and checks all 507 actual FP16 weight hashes before/after both arms; its CPU audit binds those recorded hashes and does not recompute GPU logits. Adapter/table transfer to other checkpoints requires separate validation.
The tagged repository contains the frozen protocol, full training/evaluation reproduction guide and retained evidence. Base weights and training corpora are downloaded separately.
Provenance and licenses
This release corresponds to GitHub
v0.1.0-q325-resurface,
commit be037e1a2635ad7b540e0a1d2e551febf647c346.
- Adapter SHA-256:
339334b3431027504fb8da4ce3167d63c8c2f309920cfb0076bde1f72a115cb7. - Selected calibration SHA-256:
a77bbd86ae609d8ef7cf90f10d0904cf95b4df8e059c9ea6cb2fd12886436fb8. - Full independent audit SHA-256:
ad2b3f14a075b0d60993f3579e99fc98bfbac8aa527f6a0c8d684cc9493e7d72.
New project code and the fresh adapter are distributed under GPL-3.0. The original NVIDIA model, Mamba implementation and StateQuant reference have their separate Apache-2.0 licenses. The root GPL license does not relicense those upstream assets or WikiText text. The base checkpoint is not included. See LICENSE and third-party notices for source attribution, StateQuant provenance and WikiText licensing details.
Method references: StateQuant and Resurface. The preserved StateQuant reference comes from the user-supplied archive identified in the notices; current availability of that external repository is not required to use this packaged runtime.
Model tree for EndlessChasing/Mamb2_8B_Recall_SQ3.25
Base model
nvidia/mamba2-8b-3t-4k