Commit History

Datasheet: add effective lexicon-usage accounting (corpus-intersecting slice; 9.90M < 10M under strict counting)
8534463
verified

juand-r commited on

Datasheet correction: MorphyNet used untruncated in the submission pipeline (160k cutoff is analysis/Tier-2 only)
3ac2e3c
verified

juand-r commited on

Add data statement (datasheet): cleaned strict-small corpus, budget accounting, tokenizer resources, checkpoint index
89753ed
verified

juand-r commited on

Fix train_meta.json: correct to m20-ms (was mislabeled v5pm-ms/mlm-0.5)
8416aa9
verified

juand-r commited on

Correct model card for v5: name, params (135M), recipe (Muon+MNTP@0.20), BLiMP 72.52
e2e601d
verified

juand-r commited on

Upload model.safetensors with huggingface_hub
1194e3a
verified

juand-r commited on

Upload tokenization_ourmorph.py with huggingface_hub
86dc8cc
verified

juand-r commited on

Upload our_tokenizer.py with huggingface_hub
38fdfb1
verified

juand-r commited on

Upload folder using huggingface_hub
2874b41
verified

juand-r commited on

initial commit
5c0359f
verified

juand-r commited on