Datasheet: add effective lexicon-usage accounting (corpus-intersecting slice; 9.90M < 10M under strict counting) 8534463 verified juand-r commited on 14 days ago
Datasheet correction: MorphyNet used untruncated in the submission pipeline (160k cutoff is analysis/Tier-2 only) 3ac2e3c verified juand-r commited on 14 days ago
Add data statement (datasheet): cleaned strict-small corpus, budget accounting, tokenizer resources, checkpoint index 89753ed verified juand-r commited on 14 days ago
Fix train_meta.json: correct to m20-ms (was mislabeled v5pm-ms/mlm-0.5) 8416aa9 verified juand-r commited on 15 days ago
Correct model card for v5: name, params (135M), recipe (Muon+MNTP@0.20), BLiMP 72.52 e2e601d verified juand-r commited on 16 days ago
Upload tokenization_ourmorph.py with huggingface_hub 86dc8cc verified juand-r commited on 16 days ago