plastic: lm_wikitext_l4
Canonical repository: Both checkpoints now live together at dmontgomery40/plastic. Use the text folder. This earlier release remains available for existing links.
Small BPE next-token model trained on WikiText-103 with MQAR batches; not an instruction-tuned assistant.
This is the implemented selective-recurrence plus linear delta-memory baseline with normalized readout, not the experimental nonlinear coordinate proposal or an exact reproduction of Sun et al. TTT-Linear. Outer training differentiates through the fast-memory updates. The external transactional harness is a separate runtime component.
Files and provenance
Checkpoint: 6,845,984 parameters, step 6000. Config, saved evaluation, license, and SHA-256 manifest are included; the text package also includes its BPE tokenizer. Training source revision is not recorded in this package; do not treat the compatible runtime revision as the training revision.
Compatible source: plastic at 269ecc9. Use Python 3.12+ and the dependencies pinned or declared by that checkout. This is custom PyTorch code, not a Transformers AutoModel package.
Load locally
Clone the source repository, check out the revision above, and run uv sync --extra dev. Run the example below with uv run python; it downloads this public checkpoint from Hugging Face:
from pathlib import Path
from huggingface_hub import snapshot_download
import torch
from plastic.config import ModelConfig
from plastic.model.lm import build_model
folder = Path(snapshot_download("dmontgomery40/plastic-text-6.85m"))
checkpoint = torch.load(folder / "checkpoint.pt", map_location="cpu", weights_only=True)
cfg = ModelConfig.from_dict(checkpoint["config"])
model = build_model(cfg).eval()
model.load_state_dict(checkpoint["model_state"])
from plastic.tokenizer.bpe import Tokenizer
tokenizer = Tokenizer.load(str(folder / "tokenizer.json"))
x = torch.tensor([tokenizer.encode("The history of computing", add_bos=True)])
with torch.no_grad():
predictions, state, signals = model(x)
# Pass state to the next call to continue the same sequence.
This example directly uses the model and permits fast-memory updates. It does not apply the transactional safety harness. Calibration/canaries/Fisher files and session data are not included in this portable checkpoint package. Follow the source project's calibration workflow before creating guarded sessions.
Recorded evaluation
Metric: NLL (nats/token), over 131,072 held-out tokens. Writes enabled: 3.470405519; writes disabled (beta_scale=0): 5.01505518. These are saved single-run evaluations, not freshly reproduced by this packaging step. Disabling writes leaves retention active. MQAR training exposure means recall is trained-task evidence.
These observations support task-specific memory utility. They do not establish architectural novelty, general intelligence, or adversarial robustness. The CPU loading check establishes serialization/runtime compatibility only, not GPU/MPS performance. See the source repository for experimental methods and limitations.
License
The included license permits noncommercial use and requires prior written permission for commercial use. It is not the standard MIT license.
- Downloads last month
- 22