distilled-m-teacher-init

Table 4, M, Distilled, teacher init of Towards Looped Models Done Right. Part II: Rethinking at Fixed Points: a distilled prefill path for huginn-m-learned-entropy0p01.

Loading this checkpoint requires the xLLM code. The weights are stored in xLLM's native format (BF16 Safetensors with an xLLM config.json and artifact_manifest.json). This is not a Hugging Face transformers checkpoint: AutoModel.from_pretrained cannot load it. Get the code at https://github.com/ifm-ai/xllm-loop and load it with its teacher through xllm.paper_part2.distill.load_student.

Model details

Model a d x d input map and 2 Transformer blocks of the teacher's width
Parameters 217,067,520
Teacher huginn-m-learned-entropy0p01
Initialization from the teacher: the input map is its injection at a zero state, the blocks copy its recurrent blocks
Target the teacher's pre-coda state after 5 recurrences, from its prelude output
Loss hidden-state MSE over the target energy, plus KL from the teacher's next-token distribution
Training recipe m_distill_teacher_init: 10,240 updates of 256 x 8,192 tokens, AdamW (0.9, 0.95), cosine schedule
Prefill the student's state, one teacher recurrence and the teacher's coda write the four KV banks; the teacher decodes at R = 5
Weights BF16 Safetensors: the trained FP32 weights rounded to BF16

The paper initialized this student from the teacher's FP32 weights. Retraining it from the released BF16 teacher starts from BF16-rounded copies.

Download

hf download IFM/LoopedLM-P2-huginn-m-learned-entropy0p01 --local-dir huginn-m-learned-entropy0p01
hf download IFM/LoopedLM-P2-distilled-m-teacher-init --local-dir distilled-m-teacher-init

The xLLM loader checks the directory against artifact_manifest.json: it rejects symbolic links and files the manifest does not list, apart from the .gitattributes file and the .cache/huggingface/ folder that hf download --local-dir adds. Download into a directory as above, not into the Hub cache (~/.cache/huggingface/hub), whose files are symbolic links.

Use

A distilled student is evaluated together with its teacher, which supplies the prelude, one recurrence and the coda. Download both, then run eval_paper_part2.py from the xLLM repository:

ENABLE_FLASH_ATTENTION_3=true python eval_paper_part2.py --artifact huginn-m-learned-entropy0p01 --student distilled-m-teacher-init \
    --data /path/to/eval-data/data.json --out out ppl

data.json and the evaluation inputs come from release/paper-part2/prepare-eval-data.py --output /path/to/eval-data.

The student's config.json pins its teacher's manifest_sha256, the digest recorded in the teacher's artifact_manifest.json (see Provenance); xllm.paper_part2.distill.load_student refuses any other teacher artifact, so use the teacher repository at the matching revision.

Provenance

  • Teacher manifest_sha256: 766ad24ef935ff5acb4fb37470347371d00c15dedbdcf0d99613ac37ac3bcaf6

artifact_manifest.json records the size and SHA-256 of every file in this repository.

Paper and citation

Towards Looped Models Done Right. Part II: Rethinking at Fixed Points: https://arxiv.org/abs/2610.06833

@article{huang2026fixedpoints,
  title   = {Towards Looped Models Done Right, Part II: Rethinking at Fixed Points},
  author  = {Benhao Huang and Chufan Shi and Junlin Chen and Shicheng Wen and Zhengzhong Liu and Eric Xing and Xuezhe Ma},
  journal = {arXiv preprint arXiv:2610.06833},
  year    = {2026}
}

License

The weights are released under the Apache License 2.0 (LICENSE); NOTICE records the tokenizer's attribution and how the artifact was prepared. The xLLM code is distributed under its own license.

Downloads last month
54
Safetensors
Model size
0.2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IFM/LoopedLM-P2-distilled-m-teacher-init

Finetuned
(1)
this model

Collection including IFM/LoopedLM-P2-distilled-m-teacher-init

Paper for IFM/LoopedLM-P2-distilled-m-teacher-init