File size: 4,711 Bytes
52e0242
 
8521e58
 
 
 
 
52e0242
8521e58
 
 
b4e42d9
 
 
 
 
8521e58
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b4e42d9
 
 
8521e58
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
---
license: mit
tags:
  - tokenization
  - byte-level
  - distillation
  - lexicon
---

# geolip-bytelex — the universal byte-information relational system

**Code:** [github.com/AbstractEyes/geolip-bytelex](https://github.com/AbstractEyes/geolip-bytelex) · install:
```bash
pip install "geolip-bytelex @ git+https://github.com/AbstractEyes/geolip-bytelex"
```

A **model-free translation matrix between tokenizer languages and byte
language**. Foreign tokens are not the same language as bytes — but any
real token *is* a byte string underneath, and the statistics of byte
information itself (not any particular model's internals) decide how
that token relates to byte-native units. This repo hosts the matrix:
every token of many tokenizers, projected through corpus-derived byte
relational structure, so that **any ByteLM derivative** can consume
foreign-tokenizer supervision (distillation, alignment, evaluation)
through one reusable interface.

Design rules, per the program's laws:

- **Model-free.** Nothing here depends on model weights. Internal model
  structure drifts over training; corpus byte statistics do not care.
- **Alphabet-parametric.** The alphabet (256 byte values today) is a
  schema parameter — it can grow or shrink; the system regenerates.
- **Gram-modular.** Relational views are declared, not hardcoded:
  char n-grams (1–4 exact), hashed large-n (9-gram class), word-grams
  via a separator predicate — add views as required (quadgram, wordgram
  combos, whatever the task needs).
- **Specials are out-of-alphabet.** Control tokens are tabled and
  flagged, never silently byte-expanded into text statistics.

## Artifacts

| path | contents |
|---|---|
| `vocab_<tokenizer>.jsonl` | Normalized token→bytes tables: `{"id", "hex", "text", "n_bytes", "is_special", "continuation"}`. Hex is authoritative; `text` is null when bytes aren't valid UTF-8. |
| `byte_lexicon_v1/` | The corpus statistics: exact char-gram tables (n=1..4), hashed 9-gram table, word-gram table, schema + provenance in `meta.json`. |
| `matrix_<tokenizer>.jsonl` | The full lexicon translation matrix: every non-special token's relational profile — `hmax`/`hargmax` (internal boundary-entropy maximum: where byte-language says the token divides), `pmin` (minimum internal PMI: cohesion), `word` (word-gram standing). |
| `matrix_summary.json` | Cross-tokenizer comparison: word fraction, divides/cohesive fractions, mean internal entropy. |

## Tokenizer roster (v1)

byte-BPE: gpt2 · Qwen3.8 (248,077 vocab) · DeepSeek-V3 · SmolLM2 ·
Llama-3.1 — tiktoken: cl100k_base · o200k_base — SentencePiece-BPE:
Mistral v0.3 · XLM-R — SP-unigram: T5 — WordPiece: bert-base-uncased
(lossy/uncased: flagged) — byte-identity control: ByT5.

All twelve extracted and round-trip spot-checked. Every generative
tokenizer in the roster attaches whitespace **leading**; byte-native
segmentations (measured on AlephLM) tend to build **trailing**-space
units — consumers must normalize the convention before comparing
boundaries.

## The loss primitives (what consumers do with this)

For distillation from any token-level teacher into any ByteLM student:

1. **Alignment endpoints** = corpus-entropy boundaries (successor
   branching entropy over the byte corpus), *not* any model's
   segmentation and *not* the teacher's token boundaries alone.
2. **Per-token weight** = cohesion (`pmin`, internal PMI): a token the
   byte-language considers one unit aligns as one endpoint-to-endpoint
   span.
3. **Non-cohesive tokens split** at their internal entropy maxima
   (`hargmax`) before matching — the token was two byte-units wearing
   one id, and the loss should know.

Teacher logits push forward to byte space along each token's byte
expansion (the exact-conversion direction of Phan et al., ICLR 2025);
chunk-level likelihood matching happens between co-boundaries (the ALM
family, arXiv:2503.20083), with the chunks defined by the corpus, so
the same matrix serves every teacher and every ByteLM student.

## Regeneration

The consuming/producing code is the standalone
[geolip-bytelex library](https://github.com/AbstractEyes/geolip-bytelex)
(pure-stdlib core, shared geolip namespace, 18-test suite):
build a `ByteLexicon(GramSchema(...))` over any corpus stream, `feed`
raw bytes, `save`; project any `vocab_*.jsonl` through
`ByteLexicon.profile`. New alphabet, new corpus, new gram views, new
tokenizers — same recipe.

Part of the AlephLLM / Mini-Beatrix program
([training record](https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training)).
AlephLM is the strongest current consumer of this structure — but the
matrix belongs to the bytes, not to any one model.