Fill-Mask
Transformers
Safetensors
English
tolm
babylm
strict-small
masked-language-modeling
factorized
ltg-bert
custom_code
Instructions to use miguelcsx/factorized-natural-dense with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use miguelcsx/factorized-natural-dense with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="miguelcsx/factorized-natural-dense", trust_remote_code=True)# Load model directly from transformers import AutoModelForMaskedLM model = AutoModelForMaskedLM.from_pretrained("miguelcsx/factorized-natural-dense", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
| language: | |
| - en | |
| license: other | |
| library_name: transformers | |
| pipeline_tag: fill-mask | |
| tags: | |
| - babylm | |
| - strict-small | |
| - masked-language-modeling | |
| - factorized | |
| - ltg-bert | |
| # FACTORIZED Natural Dense | |
| Selected strict-small checkpoint from the controlled FACTORIZED study. The | |
| model uses an LTG-BERT-style encoder trained with whole-word masking and an | |
| auxiliary replaced-token objective. Its 384-dimensional representation is | |
| partitioned into three 128-dimensional factor readouts. In this natural-dense | |
| control the factor readout is dormant during language-model inference. | |
| ## Model summary | |
| | Property | Value | | |
| |---|---:| | |
| | Parameters | approximately 33.3M | | |
| | Layers / hidden size | 12 / 384 | | |
| | Attention heads | 6 | | |
| | FFN size | 1,280 | | |
| | Context length | 512 | | |
| | Tokenizer | byte-level BPE, 16,384 tokens | | |
| | Training corpus | BabyLM strict-small, conservatively counted as 10M words | | |
| | Total exposure | 100M words | | |
| | Optimizer | LAMB | | |
| | Peak learning rate | 3.5e-3 | | |
| ## Usage | |
| ```python | |
| from transformers import AutoModelForMaskedLM, AutoTokenizer | |
| repo_id = "miguelcsx/factorized-natural-dense" | |
| revision = "chck_100M" | |
| tokenizer = AutoTokenizer.from_pretrained(repo_id, revision=revision) | |
| model = AutoModelForMaskedLM.from_pretrained( | |
| repo_id, | |
| revision=revision, | |
| trust_remote_code=True, | |
| ) | |
| ``` | |
| Remote code is required for the custom TOLM architecture. Review `tolm.py` | |
| before loading it in a security-sensitive environment. | |
| ## Development evaluation | |
| The selected `chck_100M` checkpoint led the FACTORIZED tournament on the | |
| pre-registered shared full-task mean among the three finalists. | |
| | Evaluation | Score | | |
| |---|---:| | |
| | BLiMP | 70.01 | | |
| | BLiMP Supplement | 60.63 | | |
| | Entity Tracking | 17.70 | | |
| | COMPS | 51.41 | | |
| | GlobalPIQA | 36.09 | | |
| These are development measurements, not a claim of an official leaderboard | |
| placement. The complete official BabyLM 2026 evaluation is being regenerated | |
| for submission. | |
| ## Checkpoint revisions | |
| The repository contains all strict-small challenge revisions: | |
| `chck_1M` through `chck_10M`, then `chck_20M` through `chck_100M`. | |
| The final submission checkpoint is `chck_100M`. | |
| ## Provenance | |
| - Final checkpoint model SHA-256: | |
| `ba89940028a902f4aa09385966b8e83bb346a0abaab6899feec1fa2d3a4b7e13`. | |
| - Evaluation backend: `mntp`. | |
| - Fast-evaluation parity maximum absolute logit delta: `0.0`. | |
| - Training and compliance details are recorded in `training_manifest.json`. | |
| ## Limitations | |
| This is a small English developmental language model trained for controlled | |
| BabyLM experiments. It is not a general-purpose assistant. The factorized | |
| readout in this control is dormant during ordinary masked-language-model | |
| inference. | |