Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

janakhpon
/
mon_tokenizer

Mon
Burmese
English
tokenizers
tokenizer
unigram
mon
burmese
myanmar
low-resource
Model card Files Files and versions
xet
Community
mon_tokenizer
4.87 MB
Ctrl+K
Ctrl+K
  • 1 contributor
History: 25 commits
janakhpon's picture
janakhpon
docs: align the status note with the card's fallback counts
0e5c937 about 24 hours ago
  • .gitattributes
    315 Bytes
    feat: restructure and upgrade to 32k vocab model (v2) 4 months ago
  • .gitignore
    784 Bytes
    chore: keep internal engineering docs off the Hub 10 days ago
  • LICENSE
    3.7 kB
    fix: stop declaring MIT over an artifact this project cannot license 5 days ago
  • NEXT_STEPS.md
    4.1 kB
    docs: align the status note with the card's fallback counts about 24 hours ago
  • README.md
    7.07 kB
    docs: split the paired dashes in the corpus-licence sentence 1 day ago
  • model_card.json
    2.72 kB
    fix: correct the single-token coverage claim to 98.74% 6 days ago
  • special_tokens_map.json
    96 Bytes
    feat: migrate the artifact from SentencePiece to tokenizers JSON 10 days ago
  • tokenizer.json
    4.86 MB
    feat: migrate the artifact from SentencePiece to tokenizers JSON 10 days ago
  • tokenizer_config.json
    240 Bytes
    feat: migrate the artifact from SentencePiece to tokenizers JSON 10 days ago