Text Generation
Transformers
Safetensors
cohere2

Set tokenizer_class to PreTrainedTokenizerFast so that transformers v5 uses tokenizer.json

#1
by peterlu02 - opened

Problem:

When the config says "tokenizer_class": "CohereTokenizer", transformers v5 loads CohereTokenizer. This class builds its own pre-tokenization rules in __init__ instead of using the rules in tokenizer.json. So under transformers v5, for the case print(type(tok).__name__, tok.tokenize("1234567")), every number is split differently from how the model was trained. We can apply exactly the same change as commit e9feb287 in the tiny-aya-l2-thinker repo. After this change, in our own tests, transformers 4.57.6 and 5.17.0 give identical results. The same change applies to tiny-aya-earth, -fire, -water and -global. I have also opened a PR on tiny-aya-earth.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment