Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
janakhpon
/
mon_tokenizer
like
0
Mon
Burmese
English
tokenizers
tokenizer
unigram
mon
burmese
myanmar
low-resource
License:
mit-code-corpus-derived-artifact
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
8fd15e7
mon_tokenizer
4.87 MB
Ctrl+K
Ctrl+K
1 contributor
History:
16 commits
janakhpon
chore: keep internal engineering docs off the Hub
8fd15e7
10 days ago
.gitattributes
Safe
315 Bytes
feat: restructure and upgrade to 32k vocab model (v2)
4 months ago
.gitignore
Safe
784 Bytes
chore: keep internal engineering docs off the Hub
10 days ago
README.md
5.76 kB
feat: migrate the artifact from SentencePiece to tokenizers JSON
11 days ago
model_card.json
Safe
2.65 kB
feat: migrate the artifact from SentencePiece to tokenizers JSON
11 days ago
special_tokens_map.json
Safe
96 Bytes
feat: migrate the artifact from SentencePiece to tokenizers JSON
11 days ago
tokenizer.json
Safe
4.86 MB
feat: migrate the artifact from SentencePiece to tokenizers JSON
11 days ago
tokenizer_config.json
Safe
240 Bytes
feat: migrate the artifact from SentencePiece to tokenizers JSON
11 days ago