Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
janakhpon
/
mon_tokenizer
like
0
Mon
Burmese
English
tokenizers
tokenizer
unigram
mon
burmese
myanmar
low-resource
License:
mit-code-corpus-derived-artifact
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
mon_tokenizer
4.87 MB
Ctrl+K
Ctrl+K
1 contributor
History:
25 commits
janakhpon
docs: align the status note with the card's fallback counts
0e5c937
about 24 hours ago
.gitattributes
Safe
315 Bytes
feat: restructure and upgrade to 32k vocab model (v2)
4 months ago
.gitignore
Safe
784 Bytes
chore: keep internal engineering docs off the Hub
10 days ago
LICENSE
3.7 kB
fix: stop declaring MIT over an artifact this project cannot license
5 days ago
NEXT_STEPS.md
4.1 kB
docs: align the status note with the card's fallback counts
about 24 hours ago
README.md
7.07 kB
docs: split the paired dashes in the corpus-licence sentence
1 day ago
model_card.json
2.72 kB
fix: correct the single-token coverage claim to 98.74%
6 days ago
special_tokens_map.json
Safe
96 Bytes
feat: migrate the artifact from SentencePiece to tokenizers JSON
10 days ago
tokenizer.json
Safe
4.86 MB
feat: migrate the artifact from SentencePiece to tokenizers JSON
10 days ago
tokenizer_config.json
Safe
240 Bytes
feat: migrate the artifact from SentencePiece to tokenizers JSON
10 days ago