ULTRON-Core Standalone Tokenizer (24,576 Vocab)

Pure-Python Byte-Level BPE Tokenizer built from scratch with zero external dependencies.

Specifications

  • Vocabulary Size: 24,576 tokens
  • Normalization: Unicode NFC Standard
  • Byte Fallback: 256-byte leaf tokens (0% Out-of-Vocabulary rate)
  • Special Tokens: <think>, </think>, <dialogue>, <turn>, <user>, <assistant>, <doc>, <para>, <code>, <math>
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support