Instructions to use AETHORIA-AI/TR-HASH-Tokenizer-32K with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AETHORIA-AI/TR-HASH-Tokenizer-32K with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("AETHORIA-AI/TR-HASH-Tokenizer-32K", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 2,063 Bytes
f1610d9 9a97c2c f1610d9 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 | ---
library_name: transformers
license: cc-by-nc-4.0
language:
- en
tags:
- tokenizer
- bpe
- byte-level
- tr-hash
---
# TR-HASH Tokenizer 32K
The 32,000-token ByteLevel BPE tokenizer used by the TR-HASH language-model
research line and the 200B-token pretraining mixture.
## Properties
- Vocabulary size: 32,000
- Model: ByteLevel BPE
- Normalization: Unicode NFC
- EOS: `</s>` (ID 0)
- PAD: `<pad>` (ID 1)
- BOS: `<s>` (ID 2)
- UNK: `<unk>` (ID 3)
- Declared model context: 2,048 tokens
## Usage
```python
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained(
"AETHORIA-AI/TR-HASH-Tokenizer-32K"
)
encoded = tokenizer("Token routing starts with tokenization.")
decoded = tokenizer.decode(encoded["input_ids"])
```
This repository contains the tokenizer only. It does not define a chat
template or a model architecture.
## Comparison with GPT `r50k_base`
`r50k_base` is used here as the approximately 50K-token GPT reference. The
comparison uses `transformers` without added special tokens for TR-HASH and
`tiktoken` for `r50k_base`.
| Encoding | Vocabulary | Tokens on fixed suite | Characters/token |
|---|---:|---:|---:|
| TR-HASH Tokenizer 32K | 32,000 | 375 | 3.712 |
| GPT `r50k_base` | 50,257 | 354 | 3.932 |
TR-HASH uses a **36.3% smaller vocabulary**. On the fixed 1,392-character
suite included in [`benchmark_r50k.py`](benchmark_r50k.py), it produces 5.9%
more tokens than `r50k_base`. At hidden size 1,024, the smaller vocabulary
removes 18,695,168 embedding parameters, or about 37.4 MB in BF16/FP16, when
input and output embeddings are tied.
This small suite covers English prose, technical text, code/JSON, mathematics,
and French. It is an illustrative, reproducible check rather than a claim
about every corpus. Compression should be measured again on the intended
training or deployment distribution.
```bash
pip install transformers tiktoken
python benchmark_r50k.py
```
## Associated dataset
[AETHORIA-AI/data-32k-200b-tokens](https://huggingface.co/datasets/AETHORIA-AI/data-32k-200b-tokens)
|