Instructions to use google/gemma-7b-it with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use google/gemma-7b-it with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="google/gemma-7b-it") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("google/gemma-7b-it") model = AutoModelForCausalLM.from_pretrained("google/gemma-7b-it", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use google/gemma-7b-it with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf google/gemma-7b-it # Run inference directly in the terminal: llama cli -hf google/gemma-7b-it
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf google/gemma-7b-it # Run inference directly in the terminal: llama cli -hf google/gemma-7b-it
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf google/gemma-7b-it # Run inference directly in the terminal: ./llama-cli -hf google/gemma-7b-it
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf google/gemma-7b-it # Run inference directly in the terminal: ./build/bin/llama-cli -hf google/gemma-7b-it
Use Docker
docker model run hf.co/google/gemma-7b-it
- LM Studio
- Jan
- vLLM
How to use google/gemma-7b-it with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "google/gemma-7b-it" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "google/gemma-7b-it", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/google/gemma-7b-it
- SGLang
How to use google/gemma-7b-it with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "google/gemma-7b-it" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "google/gemma-7b-it", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "google/gemma-7b-it" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "google/gemma-7b-it", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use google/gemma-7b-it with Ollama:
ollama run hf.co/google/gemma-7b-it
- Unsloth Desktop
- Docker Model Runner
How to use google/gemma-7b-it with Docker Model Runner:
docker model run hf.co/google/gemma-7b-it
- Lemonade
How to use google/gemma-7b-it with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull google/gemma-7b-it
Run and chat with the model
lemonade run user.gemma-7b-it-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
fix: Remediate 3736 glitch tokens in tokenizer.json
Glitch Token Remediation — Tokenizer Patch
Summary
This PR patches tokenizer.json to remediate 3736 glitch tokens identified
by embedding vector analysis. Glitch tokens are vocabulary entries whose embeddings
have collapsed to near-identical representations, causing unpredictable model behavior
when they appear in input.
No model weights are modified — this is a tokenizer-only fix.
Technique: Hybrid (merge pruning + placeholder rename)
The patch uses two complementary strategies:
1. Merge pruning (586 tokens): Removes the BPE merge rules that
produce multi-character glitch tokens. The vocab entry stays (ID preserved) but
becomes unreachable — text that would have matched is now tokenized as trained
constituent subwords instead.
2. Placeholder rename (3150 tokens): For single-character glitch tokens
(rare Unicode characters) that have no producing merge rule, the vocab string is renamed to<glitch_pruned_ID>. The original character now falls through to byte_fallback,
encoding as well-trained UTF-8 byte tokens.
Safety: Placeholder strings cannot be triggered by user input. BPE builds tokens
bottom-up from characters via merge rules only — since no merge rule produces these strings,
they are permanently unreachable. Gemma's split-digits tokenization policy provides an
additional layer of safety.
1027 merges were removed for 586 merge-pruned tokens because some glitch tokens have more than one BPE merge path, and every path must be removed.
10 flagged IDs were left unchanged because they are special tokens, byte-fallback tokens, or single ASCII characters, which the tokenizer must keep.
Before / after examples: google/gemma-7b-it
Merge prune example: bildtitel, token ID 225065
Prompt: Please repeat the following string exactly, with nothing else: "bildtitel"
| Tokens for the string | Model output | |
|---|---|---|
| Before | ['bildtitel'] → [225065] |
Sure, here is the repeated string: … ❌ |
| After | ['bild', 'titel'] → [13328, 71990] |
bildtitel … ✅ |
Placeholder rename example: ꎬ (U+A3AC), token ID 254885
Prompt: Please repeat the following string exactly, with nothing else: "ꎬ"
| Tokens for the string | Model output | |
|---|---|---|
| Before | ['ꎬ'] → [254885] |
Sure, here is the repeated string: … ❌ |
| After | ['<0xEA>', '<0x8E>', '<0xAC>'] → [451, 359, 389] |
ꎬ … ✅ |
⚠️ Note for Agentic Systems
Many remediated glitch tokens are shared across Gemma model families (Gemma 1, 2, 3, 4).
In agentic pipelines, care should be taken that these token strings are not inadvertently
injected into prompts for unpatched models. We recommend applying this patch consistently
across all Gemma models used in a pipeline.
What is preserved
- ✅
len(tokenizer)— unchanged (256,000) - ✅ All token IDs — stable, no re-indexing
- ✅ Chat template — identical to original
- ✅
tokenizer_config.json— identical to original - ✅ All special tokens and
added_tokens— unchanged - ✅ Normal text tokenization — verified via benchmarks
- ✅ Zero model weight changes
Diff summary
| Original | Patched | Delta | |
|---|---|---|---|
| Vocab size | 256,000 | 256,000 | 0 |
| Merges | 580,604 | 579,577 | −1027 |
| Vocab renamed | — | 3150 | +3150 |
Vocab renames (3150 entries)
Click to expand the first 200 renamed vocab entries
Showing the first 200 of 3150 renamed entries (renamed_vocab.tsv). Full list: https://huggingface.co/gkielian/gemma-glitch-staging/tree/main/google_gemma-7b-it
| ID | Original | Patched |
|---|---|---|
| 238396 | (U+E653) |
<glitch_pruned_238396> |
| 238562 | 雰 (U+96F0) |
<glitch_pruned_238562> |
| 239961 | 肪 (U+80AA) |
<glitch_pruned_239961> |
| 240092 | 蝴 (U+8774) |
<glitch_pruned_240092> |
| 240390 | 繍 (U+7E4D) |
<glitch_pruned_240390> |
| 240483 | 拶 (U+62F6) |
<glitch_pruned_240483> |
| 240610 | (U+E70D) |
<glitch_pruned_240610> |
| 240632 | 蜘 (U+8718) |
<glitch_pruned_240632> |
| 241191 | ܙ (U+0719) |
<glitch_pruned_241191> |
| 241373 | 琲 (U+7432) |
<glitch_pruned_241373> |
| 241602 | 曖 (U+66D6) |
<glitch_pruned_241602> |
| 241690 | 稣 (U+7A23) |
<glitch_pruned_241690> |
| 241874 | 牺 (U+727A) |
<glitch_pruned_241874> |
| 242075 | ܀ (U+0700) |
<glitch_pruned_242075> |
| 242265 | 蝙 (U+8759) |
<glitch_pruned_242265> |
| 242271 | 痤 (U+75E4) |
<glitch_pruned_242271> |
| 242301 | 赁 (U+8D41) |
<glitch_pruned_242301> |
| 242364 | (U+E5E5) |
<glitch_pruned_242364> |
| 242409 | 饪 (U+996A) |
<glitch_pruned_242409> |
| 242439 | (U+F0004) |
<glitch_pruned_242439> |
| 242588 | 螃 (U+8783) |
<glitch_pruned_242588> |
| 242594 | 諏 (U+8ACF) |
<glitch_pruned_242594> |
| 243397 | ܂ (U+0702) |
<glitch_pruned_243397> |
| 243533 | (U+F732) |
<glitch_pruned_243533> |
| 243871 | ༢ (U+0F22) |
<glitch_pruned_243871> |
| 244010 | (U+F733) |
<glitch_pruned_244010> |
| 244147 | 哐 (U+54D0) |
<glitch_pruned_244147> |
| 244366 | ༤ (U+0F24) |
<glitch_pruned_244366> |
| 244422 | (U+F739) |
<glitch_pruned_244422> |
| 244450 | ܃ (U+0703) |
<glitch_pruned_244450> |
| 244457 | ߊ (U+07CA) |
<glitch_pruned_244457> |
| 244516 | Ꮅ (U+13B5) |
<glitch_pruned_244516> |
| 244537 | Ꮓ (U+13C3) |
<glitch_pruned_244537> |
| 244548 | ཎ (U+0F4E) |
<glitch_pruned_244548> |
| 244550 | (U+F735) |
<glitch_pruned_244550> |
| 244583 | (U+F734) |
<glitch_pruned_244583> |
| 244700 | ܼ (U+073C) |
<glitch_pruned_244700> |
| 244833 | (U+F736) |
<glitch_pruned_244833> |
| 244870 | (U+F737) |
<glitch_pruned_244870> |
| 244935 | (U+F738) |
<glitch_pruned_244935> |
| 244949 | ␊ (U+240A) |
<glitch_pruned_244949> |
| 244977 | ߬ (U+07EC) |
<glitch_pruned_244977> |
| 244997 | (U+F25C) |
<glitch_pruned_244997> |
| 244998 | ༩ (U+0F29) |
<glitch_pruned_244998> |
| 245055 | ༣ (U+0F23) |
<glitch_pruned_245055> |
| 245126 | (U+F25D) |
<glitch_pruned_245126> |
| 245130 | ܶ (U+0736) |
<glitch_pruned_245130> |
| 245252 | ܤ (U+0724) |
<glitch_pruned_245252> |
| 245474 | ༔ (U+0F14) |
<glitch_pruned_245474> |
| 245484 | ܳ (U+0733) |
<glitch_pruned_245484> |
| 245501 | 咣 (U+54A3) |
<glitch_pruned_245501> |
| 245565 | ༧ (U+0F27) |
<glitch_pruned_245565> |
| 245707 | ཊ (U+0F4A) |
<glitch_pruned_245707> |
| 245810 | ౧ (U+0C67) |
<glitch_pruned_245810> |
| 245814 | ܿ (U+073F) |
<glitch_pruned_245814> |
| 245817 | (U+71706) |
<glitch_pruned_245817> |
| 245828 | ܰ (U+0730) |
<glitch_pruned_245828> |
| 245837 | (U+E632) |
<glitch_pruned_245837> |
| 245844 | (U+E63B) |
<glitch_pruned_245844> |
| 245859 | ཌ (U+0F4C) |
<glitch_pruned_245859> |
| 246006 | ༦ (U+0F26) |
<glitch_pruned_246006> |
| 246059 | ཿ (U+0F7F) |
<glitch_pruned_246059> |
| 246069 | 婳 (U+5A73) |
<glitch_pruned_246069> |
| 246106 | ⬍ (U+2B0D) |
<glitch_pruned_246106> |
| 246236 | ߫ (U+07EB) |
<glitch_pruned_246236> |
| 246291 | (U+F644) |
<glitch_pruned_246291> |
| 246325 | ߌ (U+07CC) |
<glitch_pruned_246325> |
| 246342 | ܲ (U+0732) |
<glitch_pruned_246342> |
| 246466 | ܺ (U+073A) |
<glitch_pruned_246466> |
| 246479 | 嗬 (U+55EC) |
<glitch_pruned_246479> |
| 246547 | ཋ (U+0F4B) |
<glitch_pruned_246547> |
| 246571 | ߟ (U+07DF) |
<glitch_pruned_246571> |
| 246588 | ྜ (U+0F9C) |
<glitch_pruned_246588> |
| 246620 | ୦ (U+0B66) |
<glitch_pruned_246620> |
| 246657 | ߲ (U+07F2) |
<glitch_pruned_246657> |
| 246693 | 靣 (U+9763) |
<glitch_pruned_246693> |
| 246712 | ୮ (U+0B6E) |
<glitch_pruned_246712> |
| 246760 | ୫ (U+0B6B) |
<glitch_pruned_246760> |
| 246843 | ༨ (U+0F28) |
<glitch_pruned_246843> |
| 246876 | ༈ (U+0F08) |
<glitch_pruned_246876> |
| 246884 | ୪ (U+0B6A) |
<glitch_pruned_246884> |
| 246905 | ⚊ (U+268A) |
<glitch_pruned_246905> |
| 247024 | 鱧 (U+9C67) |
<glitch_pruned_247024> |
| 247025 | (U+F6DC) |
<glitch_pruned_247025> |
| 247041 | (U+F645) |
<glitch_pruned_247041> |
| 247125 | ߘ (U+07D8) |
<glitch_pruned_247125> |
| 247190 | ܽ (U+073D) |
<glitch_pruned_247190> |
| 247215 | (U+E901) |
<glitch_pruned_247215> |
| 247283 | ႏ (U+108F) |
<glitch_pruned_247283> |
| 247313 | Ꮵ (U+13E5) |
<glitch_pruned_247313> |
| 247377 | ཥ (U+0F65) |
<glitch_pruned_247377> |
| 247380 | 囘 (U+56D8) |
<glitch_pruned_247380> |
| 247408 | (U+F232) |
<glitch_pruned_247408> |
| 247409 | ༑ (U+0F11) |
<glitch_pruned_247409> |
| 247426 | Ꮝ (U+13CD) |
<glitch_pruned_247426> |
| 247445 | ꩻ (U+AA7B) |
<glitch_pruned_247445> |
| 247458 | 曏 (U+66CF) |
<glitch_pruned_247458> |
| 247524 | 玧 (U+73A7) |
<glitch_pruned_247524> |
| 247572 | 鲳 (U+9CB3) |
<glitch_pruned_247572> |
| 247586 | Ꭻ (U+13AB) |
<glitch_pruned_247586> |
| 247635 | ܆ (U+0706) |
<glitch_pruned_247635> |
| 247641 | ܇ (U+0707) |
<glitch_pruned_247641> |
| 247659 | (U+E145) |
<glitch_pruned_247659> |
| 247685 | 胗 (U+80D7) |
<glitch_pruned_247685> |
| 247739 | Ꮫ (U+13DB) |
<glitch_pruned_247739> |
| 247780 | ྞ (U+0F9E) |
<glitch_pruned_247780> |
| 247798 | ῎ (U+1FCE) |
<glitch_pruned_247798> |
| 247805 | ޑ (U+0791) |
<glitch_pruned_247805> |
| 247822 | ⃜ (U+20DC) |
<glitch_pruned_247822> |
| 247829 | ޮ (U+07AE) |
<glitch_pruned_247829> |
| 247889 | ߎ (U+07CE) |
<glitch_pruned_247889> |
| 247893 | ﮐ (U+FB90) |
<glitch_pruned_247893> |
| 247913 | ᩠ (U+1A60) |
<glitch_pruned_247913> |
| 247969 | ߛ (U+07DB) |
<glitch_pruned_247969> |
| 247989 | 嘁 (U+5601) |
<glitch_pruned_247989> |
| 248001 | ޓ (U+0793) |
<glitch_pruned_248001> |
| 248028 | (U+F643) |
<glitch_pruned_248028> |
| 248079 | ੬ (U+0A6C) |
<glitch_pruned_248079> |
| 248124 | ߍ (U+07CD) |
<glitch_pruned_248124> |
| 248131 | 豇 (U+8C47) |
<glitch_pruned_248131> |
| 248137 | ߡ (U+07E1) |
<glitch_pruned_248137> |
| 248145 | 茼 (U+833C) |
<glitch_pruned_248145> |
| 248178 | Ꮴ (U+13E4) |
<glitch_pruned_248178> |
| 248199 | (U+F646) |
<glitch_pruned_248199> |
| 248221 | ꪳ (U+AAB3) |
<glitch_pruned_248221> |
| 248224 | ܸ (U+0738) |
<glitch_pruned_248224> |
| 248237 | ୂ (U+0B42) |
<glitch_pruned_248237> |
| 248245 | ށ (U+0781) |
<glitch_pruned_248245> |
| 248246 | ፥ (U+1365) |
<glitch_pruned_248246> |
| 248263 | ߓ (U+07D3) |
<glitch_pruned_248263> |
| 248269 | 莜 (U+839C) |
<glitch_pruned_248269> |
| 248309 | 〬 (U+302C) |
<glitch_pruned_248309> |
| 248331 | 鰆 (U+9C06) |
<glitch_pruned_248331> |
| 248337 | (U+F21D) |
<glitch_pruned_248337> |
| 248360 | Ꮕ (U+13C5) |
<glitch_pruned_248360> |
| 248384 | ྛ (U+0F9B) |
<glitch_pruned_248384> |
| 248403 | ⤥ (U+2925) |
<glitch_pruned_248403> |
| 248433 | ୃ (U+0B43) |
<glitch_pruned_248433> |
| 248440 | 炝 (U+709D) |
<glitch_pruned_248440> |
| 248449 | ⤦ (U+2926) |
<glitch_pruned_248449> |
| 248453 | 箬 (U+7BAC) |
<glitch_pruned_248453> |
| 248464 | 𐌰 (U+10330) |
<glitch_pruned_248464> |
| 248470 | ߙ (U+07D9) |
<glitch_pruned_248470> |
| 248471 | ⮑ (U+2B91) |
<glitch_pruned_248471> |
| 248493 | (U+F111) |
<glitch_pruned_248493> |
| 248517 | (U+F63A) |
<glitch_pruned_248517> |
| 248520 | 恹 (U+6079) |
<glitch_pruned_248520> |
| 248580 | ߐ (U+07D0) |
<glitch_pruned_248580> |
| 248593 | Ꭴ (U+13A4) |
<glitch_pruned_248593> |
| 248604 | 茭 (U+832D) |
<glitch_pruned_248604> |
| 248619 | ྶ (U+0FB6) |
<glitch_pruned_248619> |
| 248649 | ௮ (U+0BEE) |
<glitch_pruned_248649> |
| 248650 | ၚ (U+105A) |
<glitch_pruned_248650> |
| 248660 | (U+E140) |
<glitch_pruned_248660> |
| 248677 | Ꮻ (U+13EB) |
<glitch_pruned_248677> |
| 248686 | (U+E042) |
<glitch_pruned_248686> |
| 248688 | (U+F080) |
<glitch_pruned_248688> |
| 248691 | ྰ (U+0FB0) |
<glitch_pruned_248691> |
| 248709 | 婠 (U+5A60) |
<glitch_pruned_248709> |
| 248724 | (U+F64C) |
<glitch_pruned_248724> |
| 248729 | ௨ (U+0BE8) |
<glitch_pruned_248729> |
| 248732 | Ꮚ (U+13CA) |
<glitch_pruned_248732> |
| 248745 | 莳 (U+83B3) |
<glitch_pruned_248745> |
| 248764 | ੯ (U+0A6F) |
<glitch_pruned_248764> |
| 248771 | (U+F647) |
<glitch_pruned_248771> |
| 248798 | (U+E3D0) |
<glitch_pruned_248798> |
| 248820 | Ꮣ (U+13D3) |
<glitch_pruned_248820> |
| 248849 | 妘 (U+5998) |
<glitch_pruned_248849> |
| 248865 | 噘 (U+5658) |
<glitch_pruned_248865> |
| 248871 | (U+F09D) |
<glitch_pruned_248871> |
| 248896 | 眎 (U+770E) |
<glitch_pruned_248896> |
| 248911 | (U+E5F1) |
<glitch_pruned_248911> |
| 248922 | 鸪 (U+9E2A) |
<glitch_pruned_248922> |
| 248925 | ੮ (U+0A6E) |
<glitch_pruned_248925> |
| 248926 | ྃ (U+0F83) |
<glitch_pruned_248926> |
| 248927 | Ꮄ (U+13B4) |
<glitch_pruned_248927> |
| 248932 | 绡 (U+7EE1) |
<glitch_pruned_248932> |
| 248959 | ㈵ (U+3235) |
<glitch_pruned_248959> |
| 248977 | ౮ (U+0C6E) |
<glitch_pruned_248977> |
| 249000 | 蛏 (U+86CF) |
<glitch_pruned_249000> |
| 249032 | 鹧 (U+9E67) |
<glitch_pruned_249032> |
| 249050 | ౨ (U+0C68) |
<glitch_pruned_249050> |
| 249068 | ㄗ (U+3117) |
<glitch_pruned_249068> |
| 249141 | ୯ (U+0B6F) |
<glitch_pruned_249141> |
| 249159 | 傩 (U+50A9) |
<glitch_pruned_249159> |
| 249193 | 膑 (U+8191) |
<glitch_pruned_249193> |
| 249205 | ⁜ (U+205C) |
<glitch_pruned_249205> |
| 249208 | 怏 (U+600F) |
<glitch_pruned_249208> |
| 249210 | 礽 (U+793D) |
<glitch_pruned_249210> |
| 249213 | 蝰 (U+8770) |
<glitch_pruned_249213> |
| 249224 | 蚬 (U+86AC) |
<glitch_pruned_249224> |
| 249228 | (U+E902) |
<glitch_pruned_249228> |
| 249232 | ௩ (U+0BE9) |
<glitch_pruned_249232> |
| 249245 | ޯ (U+07AF) |
<glitch_pruned_249245> |
| 249260 | (U+F639) |
<glitch_pruned_249260> |
| 249266 | 〡 (U+3021) |
<glitch_pruned_249266> |
| 249283 | 闿 (U+95FF) |
<glitch_pruned_249283> |
| 249285 | 鲅 (U+9C85) |
<glitch_pruned_249285> |
| 249290 | ௫ (U+0BEB) |
<glitch_pruned_249290> |
| 249293 | Ꮟ (U+13CF) |
<glitch_pruned_249293> |
Removed merges (1027)
Click to expand the first 200 of 1027 removed merges
Showing the first 200 of 1027 removed merges (removed_merges.txt). Full list: https://huggingface.co/gkielian/gemma-glitch-staging/tree/main/google_gemma-7b-it
["">", " "] → ">
["">", "↑↑↑</"] → ">↑↑↑</
["">", "😂"] → ">😂
["$_", "?"] → $_?
[")$", "_,"] → )$_,
[")", "$_,"] → )$_,
[")$_", ","] → )$_,
[")$_", "."] → )$_.
[")$", "_."] → )$_.
[")", "$_."] → )$_.
[")", ")$_"] → ))$_
["))", "$_"] → ))$_
["))$", "_"] → ))$_
["=", " "] → =
[">", "\<^"] → >\<^
[">\", "<^"] → >\<^
[">\<", "^"] → >\<^
["Adv", "Ex"] → AdvEx
["Answer", "Step"] → AnswerStep
["At", "bildēt"] → Atbildēt
["Bak", "grunns"] → Bakgrunns
["Curtir", "Curtir"] → CurtirCurtir
["Di", "berdayakan"] → Diberdayakan
["Di", "wed"] → Diwed
["Diwedd", "ar"] → Diweddar
["Diwed", "dar"] → Diweddar
["Dona", "tivos"] → Donativos
["Don", "ativos"] → Donativos
["English", "Choose"] → EnglishChoose
["GE", "BUR"] → GEBUR
["Gilla", "Gilla"] → GillaGilla
["IB", "LIO"] → IBLIO
["ICTO", "GRAM"] → ICTOGRAM
["Krank", "heitsbild"] → Krankheitsbild
["Line", "VCR"] → LineVCR
["Nd", "Ex"] → NdEx
["Num", "erade"] → Numerade
["Numer", "ade"] → Numerade
["OUT", "UBE"] → OUTUBE
["Parameter", "Limits"] → ParameterLimits
["ParameterLimits", "V"] → ParameterLimitsV
["Pg", "ina"] → Pgina
["P", "gina"] → Pgina
["Privat", "patient"] → Privatpatient
["R", "uju"] → Ruju
["Ruj", "u"] → Ruju
["Ru", "ju"] → Ruju
["Spo", "lja"] → Spolja
["Sp", "olja"] → Spolja
["VERTIS", "EMENT"] → VERTISEMENT
["Väl", "isling"] → Välisling
["Window", "LineVCR"] → WindowLineVCR
["\<", "^"] → \<^
["\", "<^"] → \<^
["^(@)", "$_"] → ^(@)$_
["^(@", ")$_"] → ^(@)$_
["acter", "ísticas"] → acterísticas
["actéris", "ti"] → actéristi
["acté", "risti"] → actéristi
["aime", "J"] → aimeJ
["alak", "ip"] → alakip
["ala", "kip"] → alakip
["archiw", "izowane"] → archiwizowane
["att", "utto"] → attutto
["attu", "tto"] → attutto
["bildschirm", "foto"] → bildschirmfoto
["bild", "titel"] → bildtitel
["bl", "uza"] → bluza
["blu", "za"] → bluza
["bon", "eca"] → boneca
["bone", "ca"] → boneca
["bra", "kk"] → brakk
["brak", "k"] → brakk
["br", "akk"] → brakk
["brin", "co"] → brinco
["br", "inco"] → brinco
["can", "eca"] → caneca
["cane", "ca"] → caneca
["capace", "te"] → capacete
["capa", "cete"] → capacete
["capac", "ete"] → capacete
["chin", "elo"] → chinelo
["chine", "lo"] → chinelo
["cit", "roën"] → citroën
["col", "gante"] → colgante
["d", "oudoune"] → doudoune
["dá", "ms"] → dáms
["dáms", "ké"] → dámské
["emen", "tara"] → ementara
["ement", "ara"] → ementara
["enab", "log"] → enablog
["ena", "blog"] → enablog
["e", "ſ"] → eſ
["fel", "pa"] → felpa
["fieur", "s"] → fieurs
["fie", "urs"] → fieurs
["fi", "eurs"] → fieurs
["geno", "digd"] → genodigd
["g", "eſ"] → geſ
["ge", "ſ"] → geſ
["ha", "ikusbot"] → haikusbot
["idler", "tid"] → idlertid
["ielle", "icht"] → ielleicht
["iel", "leicht"] → ielleicht
["ient", "ras"] → ientras
["ien", "tras"] → ientras
["i", "entras"] → ientras
["iffa", "nce"] → iffance
["iff", "ance"] → iffance
["ikus", "bot"] → ikusbot
["i", "neſs"] → ineſs
["ine", "ſs"] → ineſs
["isGrid", "AdvEx"] → isGridAdvEx
["is", "Ora"] → isOra
["isOra", "ColElement"] → isOraColElement
["itsub", "ishi"] → itsubishi
["i", "ſche"] → iſche
["iſ", "che"] → iſche
["iſ", "chen"] → iſchen
["i", "ſchen"] → iſchen
["iſche", "n"] → iſchen
["i", "ſe"] → iſe
["iſ", "e"] → iſe
["i", "ſen"] → iſen
["iſe", "n"] → iſen
["iſ", "en"] → iſen
["iſ", "h"] → iſh
["i", "ſh"] → iſh
["i", "ſten"] → iſten
["iſ", "ten"] → iſten
["iſt", "en"] → iſten
["iſ", "ter"] → iſter
["i", "ſter"] → iſter
["iſt", "er"] → iſter
["ja", "queta"] → jaqueta
["ke", "meja"] → kemeja
["kos", "zulka"] → koszulka
["kur", "tka"] → kurtka
["kurt", "ka"] → kurtka
["l", "brakk"] → lbrakk
["le", "ſs"] → leſs
["lina", "wan"] → linawan
["lin", "awan"] → linawan
["lla", "vero"] → llavero
["llave", "ro"] → llavero
["l", "pgp"] → lpgp
["lp", "gp"] → lpgp
["lx", "task"] → lxtask
["lxt", "ask"] → lxtask
["maca", "cão"] → macacão
["mac", "acão"] → macacão
["maj", "ánló"] → majánló
["mak", "eat"] → makeat
["make", "at"] → makeat
["mb", "gg"] → mbgg
["mbg", "g"] → mbgg
["m", "pagne"] → mpagne
["mp", "agne"] → mpagne
["mpa", "gne"] → mpagne
["ni", "ſſe"] → niſſe
["o", "uſ"] → ouſ
["ou", "ſ"] → ouſ
["pe", "cabe"] → pecabe
["pec", "abe"] → pecabe
["pend", "entif"] → pendentif
["pendenti", "f"] → pendentif
["ping", "ente"] → pingente
["pin", "gente"] → pingente
["posts", "leuth"] → postsleuth
["puls", "eira"] → pulseira
["pulse", "ira"] → pulseira
["r", "brakk"] → rbrakk
["rugu", "ay"] → ruguay
["sand", "alia"] → sandalia
["sandal", "ia"] → sandalia
["sand", "ália"] → sandália
["save", "videobot"] → savevideobot
["savevideo", "bot"] → savevideobot
["scher", "mata"] → schermata
["spod", "nie"] → spodnie
["sprzed", "am"] → sprzedam
["sud", "adera"] → sudadera
["s", "udadera"] → sudadera
["suki", "enka"] → sukienka
["suk", "ienka"] → sukienka
["sí", "ða"] → síða
["tama", "ris"] → tamaris
["tam", "aris"] → tamaris
["ta", "maris"] → tamaris
["ti", "érrez"] → tiérrez
["trauer", "anzeige"] → traueranzeige
["utter", "stock"] → utterstock
["utters", "tock"] → utterstock
["u", "ſe"] → uſe
["uſ", "e"] → uſe
["ver", "wijs"] → verwijs
["vid", "axl"] → vidaxl
["vida", "xl"] → vidaxl
["vide", "obot"] → videobot
["video", "bot"] → videobot
Files changed
tokenizer.json— BPE merge rules pruned, 3150 vocab entries renamed
Algorithm Details
For full details on the glitch token collection algorithm and remediation techniques,
reach out to gkielian.
Related
This is part of a series of tokenizer fixes across all Gemma model repositories
(Gemma 1, 2, 3, 4, and MedGemma — 18 models total).