File size: 1,603 Bytes
38c3542 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 | ---
license: apache-2.0
base_model: allenai/Olmo-3-7B-Instruct
tags:
- tokenizer
---
# Olmo-3-Instruct tokenizer (`eos_token` fixed to `<|im_end|>`)
Tokenizer-only fork of the Olmo 3 Instruct family (e.g.
[`allenai/Olmo-3-7B-Instruct`](https://huggingface.co/allenai/Olmo-3-7B-Instruct)) —
should work for any size sharing the same Instruct tokenizer/chat template.
No model weights here — use this repo for the tokenizer/chat template only,
alongside the original repo's weights.
## What changed
`eos_token` is set to `<|im_end|>` (token id `100265`) instead of the
upstream default `<|endoftext|>` (token id `100257`). Everything else
(vocab, `chat_template.jinja`, special tokens) is untouched.
## Why
Olmo 3's chat template only emits the literal `eos_token` after the final
assistant turn *when that turn is the last message in the list*:
```jinja
{%- if not loop.last -%}{{- '<|im_end|>' + '\n' -}}
{%- else -%}{{- eos_token -}}{%- endif -%}
```
Every other turn boundary closes with `<|im_end|>` unconditionally.
Tooling that renders a partial trajectory (e.g. to find the templated
prefix before the last assistant turn, without that turn present) never
sees the real `eos_token` and errors — even though `<|im_end|>` (id
`100265`) is already registered as a valid stop token in the model's own
`generation_config.json` (`eos_token_id: [100265, 100257]`), and is a real,
official special token (`special: true`) in the model's own vocabulary.
This fork just makes `tokenizer.eos_token` point at the one that actually
appears at every turn boundary, sidestepping the conditional.
|