MP1 โ€” Muon + self-distillation + causal memory

Frozen DASE7506 MP1 predictor. Full-test BPB: 1.4426896653708432, identical in three CPU FP32 runs with the original course scorer. Validation BPB: 1.4299011468116516.

  • Code and reproduction instructions: KokoNAa/7506_mp1 at the verified release commit.
  • Exact code revision: a8959d19b67d17086ae23885881aeca75cb1f40d.
  • Fixed vocabulary: course BPE-2048; independent causal windows of 256 targets.
  • Neural model: four layers, width 384, four heads, RoPE, GELU, 7,885,056 trainable parameters.
  • Selected neural checkpoint: step 23,700 of a 24,000-step own-trajectory run, with fixed causal cache and train-derived n-gram statistics.
  • Predictor SHA256: 2806c906b9d32f1fa01ce56aec16267d6343b348f2fc03490899da800738cc78.

Files

File Role
model.pt Complete inference checkpoint, including the training-derived memory tables.
baseline.pt Original 1,200-step course baseline for resource comparisons.
training/teacher-1.pt, training/teacher-2.pt Optional frozen, course-trained teachers for reproducing student training. They are not loaded during evaluation.
checksums.json, SHA256SUMS Exact file sizes, roles and SHA256 checksums.
test-results.json Preserved full-test scoring summary.

This is a custom PyTorch checkpoint used with the matching repository code. It is not a Transformers AutoModel bundle and requires no retraining for evaluation. In the code repository, run:

python -m pip install -r requirements.txt -r requirements-download.txt
python scripts/download_weights.py
python scripts/verify_release.py
cd code
python evaluate.py --checkpoint ../weights/model.pt --device cpu --precision fp32 --threads 4 --split test

The code release pins a Hugging Face commit and verifies the checkpoint hash. The only weight file required for scoring the candidate is model.pt; optional baseline and teacher artifacts are separate predictors and are excluded from candidate inference-asset accounting.

Measured resources

Three predeclared interleaved baseline/candidate full-test pairs on the same host:

  • Median CPU scoring-time ratio: 4.0429328636ร— baseline, below 5ร—.
  • Peak evaluation RAM: 2.1739997864 GiB, below 4 GiB.
  • Frozen uncompressed inference assets, including the matching source and tokenizer: 53.1046886444 MiB, below 64 MiB.

The baseline score is 2.101258027231201 BPB. Each pass covers 428,405 targets and 1,292,013 UTF-8 bytes. Timings depend on hardware. The original scoring-time summary contains a historical local_backup_pending field; the later local audit and complete hash-verified backup were completed before publication.

Training, attribution and limitations

All neural weights and n-gram statistics derive from the supplied course training text. The student starts from random initialization; teachers are existing self-trained course models. The full student trajectory processes 196,608,000 targets and the same number of forward targets per teacher. Teacher training ancestry adds 294,912,000 targets. Teacher optimization does not continue during distillation.

Temporary caches read only already observed prefixes within the current independent window. There is no persistent evaluation-prefix state, cached validation/test-answer table or evaluation network access.

Earlier model test results had been observed before subsequent development. That chronology is disclosed in the code repository. The present checkpoint and method were frozen before its own three test repetitions, with no tuning during those repetitions. Instructor acceptance of the earlier post-test development remains unconfirmed. Uploading these artifacts does not submit a course leaderboard entry.

OpenAI Codex/ChatGPT substantially assisted implementation, experiment operations, analysis and documentation. The course starter and cited methods are acknowledged in the code repository. Dataset attribution and notices are retained there; no new blanket license is assigned to the supplied classroom code or this repository by this model card.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support