SecludedCorner's picture
model card
927fe70 verified
|
Raw
History Blame Contribute Delete
2.28 kB
---
language: en
license: cc-by-4.0
library_name: transformers
pipeline_tag: text-generation
inference: false
tags:
- babylm
- babylm-2026
- strict-small
- ablation
- custom_code
---
# bind1 — trained-ablation checkpoints (BabyLM 2026 Strict-Small)
Companion repo to the entry
[`SecludedCorner/bind1-babylm2026-strict-small`](https://huggingface.co/SecludedCorner/bind1-babylm2026-strict-small):
every **retrained** ablation family behind the papers' claims, as loadable checkpoints
(one branch each, `trust_remote_code`). Eval-time ablations (severed edge, forced-$T$) need no
weights of their own — they are config-only clones of the entry; scripts in the
[code repo](https://github.com/SecludedCorner/bind1-babylm2026).
## Branches
| Branch | What it is | Headline number |
|---|---|---|
| `tt1_seed0` … `tt1_seed9` | identical architecture trained **single-pass** ($T{=}1$), ten seeds | entity tracking never forms: 17.4±2.4 (chance ≈ 20.0) in 10/10, grammar healthy (BLiMP 65.4±0.9) — **training-time iteration is the scaffold** |
| `novg_seed0` … `novg_seed3` | trained with the verdict-to-trust edge frozen at zero, four seeds | single-pass write fails to form in 3/4 (19–30) and formed anyway in one (39.7) — the edge raises the odds, **not strictly necessary** |
Loading any branch:
```python
model = AutoModelForCausalLM.from_pretrained(
"SecludedCorner/bind1-babylm2026-ablations", revision="tt1_seed0", trust_remote_code=True)
```
Per-item evaluation outputs for all of these live in the
[eval-artifacts dataset](https://huggingface.co/datasets/SecludedCorner/bind1-babylm2026-eval-artifacts).
Internal ids: `tt1_seedN` = `bind1_tt1_sN`; `novg_seedN` = `bind1_tt3_sN_novg`
(physical `loop2_novg(_sN)`).
## Honest note
These are the ablations that *disagreed with us* as often as they agreed: the ten-seed test
falsified the "single-pass training might suffice" reading, and seed 3 of the frozen-edge
family falsified "the edge is strictly necessary." Released so the disagreements are
verifiable too.
## Citation
Yulin Yang (ORCID 0009-0007-4827-8449). *A Microkernel Language Model: Reasoning as
Mutually-Supporting Aggregates, and Why Its Parts Must Be Judged Together.* BabyLM Challenge
2026 (Strict-Small track) submission.