Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
juand-r
/
morpheus-10M-v5
like
0
Safetensors
English
llama
babylm
babylm-2026
strict-small
morphological-tokenizer
License:
mit
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
morpheus-10M-v5
Ctrl+K
Ctrl+K
1 contributor
History:
10 commits
juand-r
Datasheet: add effective lexicon-usage accounting (corpus-intersecting slice; 9.90M < 10M under strict counting)
8534463
verified
13 days ago
data
Upload folder using huggingface_hub
15 days ago
resources
Upload folder using huggingface_hub
15 days ago
.gitattributes
Safe
1.65 kB
Upload folder using huggingface_hub
15 days ago
README.md
4.87 kB
Datasheet: add effective lexicon-usage accounting (corpus-intersecting slice; 9.90M < 10M under strict counting)
13 days ago
config.json
Safe
654 Bytes
Upload folder using huggingface_hub
15 days ago
generation_config.json
Safe
111 Bytes
Upload folder using huggingface_hub
15 days ago
model.safetensors
540 MB
xet
Upload model.safetensors with huggingface_hub
15 days ago
morph_tokenizer.py
Safe
36.9 kB
Upload folder using huggingface_hub
15 days ago
our_tokenizer.py
Safe
11.5 kB
Upload our_tokenizer.py with huggingface_hub
15 days ago
our_vocab.json
Safe
545 kB
Upload folder using huggingface_hub
15 days ago
resources_resolved.json
Safe
16 MB
xet
Upload folder using huggingface_hub
15 days ago
tokenization_ourmorph.py
Safe
6.62 kB
Upload tokenization_ourmorph.py with huggingface_hub
15 days ago
tokenizer_config.json
Safe
205 Bytes
Upload folder using huggingface_hub
15 days ago
train_meta.json
Safe
354 Bytes
Fix train_meta.json: correct to m20-ms (was mislabeled v5pm-ms/mlm-0.5)
14 days ago