fix: Remediate 32 glitch tokens in tokenizer.json

#85

Glitch Token Remediation — Tokenizer Patch

Summary

This PR patches tokenizer.json to remediate 32 glitch tokens identified
by embedding vector analysis. Glitch tokens are vocabulary entries whose embeddings
have collapsed to near-identical representations, causing unpredictable model behavior
when they appear in input.

No model weights are modified — this is a tokenizer-only fix.

Technique: Hybrid (merge pruning + placeholder rename)

The patch uses two complementary strategies:

1. Merge pruning (8 tokens): Removes the BPE merge rules that
produce multi-character glitch tokens. The vocab entry stays (ID preserved) but
becomes unreachable — text that would have matched is now tokenized as trained
constituent subwords instead.

2. Placeholder rename (24 tokens): For single-character glitch tokens
(rare Unicode characters) that have no producing merge rule, the vocab string is renamed to
<glitch_pruned_ID>. The original character now falls through to byte_fallback,
encoding as well-trained UTF-8 byte tokens.

Safety: Placeholder strings cannot be triggered by user input. BPE builds tokens
bottom-up from characters via merge rules only — since no merge rule produces these strings,
they are permanently unreachable. Gemma's split-digits tokenization policy provides an
additional layer of safety.

10 merges were removed for 8 merge-pruned tokens because some glitch tokens have more than one BPE merge path, and every path must be removed.

Before / after examples: google/gemma-2-2b-it

Merge prune example: bildtitel, token ID 225065

Prompt: Please repeat the following string exactly, with nothing else: "bildtitel"

Tokens for the string Model output
Before ['bildtitel'] → [225065] " ❌
After ['bild', 'titel'] → [13328, 71990] bildtitel ✅

Placeholder rename example: 𑄣 (U+11123), token ID 253904

Prompt: Please repeat the following string exactly, with nothing else: "𑄣"

Tokens for the string Model output
Before ['𑄣'] → [253904] " ❌
After ['<0xF0>', '<0x91>', '<0x84>', '<0xA3>'] → [457, 362, 349, 380] 𑄣 ✅

⚠️ Note for Agentic Systems

Many remediated glitch tokens are shared across Gemma model families (Gemma 1, 2, 3, 4).
In agentic pipelines, care should be taken that these token strings are not inadvertently
injected into prompts for unpatched models. We recommend applying this patch consistently
across all Gemma models used in a pipeline.

What is preserved

  • ✅ len(tokenizer) — unchanged (256,000)
  • ✅ All token IDs — stable, no re-indexing
  • ✅ Chat template — identical to original
  • ✅ tokenizer_config.json — identical to original
  • ✅ All special tokens and added_tokens — unchanged
  • ✅ Normal text tokenization — verified via benchmarks
  • ✅ Zero model weight changes

Diff summary

Original Patched Delta
Vocab size 256,000 256,000 0
Merges 580,604 580,594 −10
Vocab renamed — 24 +24

Vocab renames (24 entries)

Click to expand all 24 renamed vocab entries
ID Original Patched
251496 ஥ (U+0BA5) <glitch_pruned_251496>
251497 ஦ (U+0BA6) <glitch_pruned_251497>
251632 𑄨 (U+11128) <glitch_pruned_251632>
251921 ஧ (U+0BA7) <glitch_pruned_251921>
252083 (U+F565) <glitch_pruned_252083>
252372 𑄢 (U+11122) <glitch_pruned_252372>
252915 (U+F3F5) <glitch_pruned_252915>
253103 𑄚 (U+1111A) <glitch_pruned_253103>
253441 (U+E984) <glitch_pruned_253441>
253613 󠁁 (U+E0041) <glitch_pruned_253613>
253841 𑄬 (U+1112C) <glitch_pruned_253841>
253904 𑄣 (U+11123) <glitch_pruned_253904>
254071 (U+EF5A) <glitch_pruned_254071>
254455 (U+ED90) <glitch_pruned_254455>
254456 (U+EFA6) <glitch_pruned_254456>
254686 𑄮 (U+1112E) <glitch_pruned_254686>
255122 (U+F540) <glitch_pruned_255122>
255123 𑄥 (U+11125) <glitch_pruned_255123>
255124 𑄪 (U+1112A) <glitch_pruned_255124>
255517 𑄝 (U+1111D) <glitch_pruned_255517>
255645 (U+EF0E) <glitch_pruned_255645>
255663 󰈻 (U+F023B) <glitch_pruned_255663>
255806 𑄠 (U+11120) <glitch_pruned_255806>
255849 ⏔ (U+23D4) <glitch_pruned_255849>

Removed merges (10)

Click to expand the full list of 10 removed merges
["bild", "titel"]  →  bildtitel
["‌آم", "باردا"]  →  ‌آمباردا
["▁Accepted", "Loading"]  →  ▁AcceptedLoading
["▁Die", "ſe"]  →  ▁Dieſe
["▁coach", "Try"]  →  ▁coachTry
["▁ſ", "elb"]  →  ▁ſelb
["▁ſe", "lb"]  →  ▁ſelb
["▁ſel", "b"]  →  ▁ſelb
["▁パン", "チラ"]  →  ▁パンチラ
["▁メ", "ンテナ"]  →  ▁メンテナ

Files changed

  • tokenizer.json — BPE merge rules pruned, 24 vocab entries renamed

Algorithm Details

For full details on the glitch token collection algorithm and remediation techniques,
reach out to gkielian.

Related

This is part of a series of tokenizer fixes across all Gemma model repositories
(Gemma 1, 2, 3, 4, and MedGemma — 18 models total).

cc @dougreid @ssmoot

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment