Polish Morphological BPE Tokenizer
A morphologically-aware BPE (Byte-Pair Encoding) tokenizer for Polish that respects morpheme boundaries during tokenization.
Key Features
- Morpheme-aware: Prevents merges across morphological boundaries (prefixes, suffixes, case endings)
- SentencePiece-style: Uses
▁(U+2581) to mark word boundaries for perfect reconstruction - POS-aware: Uses spaCy for part-of-speech tagging to apply correct morpheme rules per word type
- Vocabulary: 30,000 tokens trained on Polish Wikipedia
Installation
# Clone the repository
git clone https://huggingface.co/rafal-adamczyk/polish-morphological-tokenizer
cd polish-morphological-tokenizer
# Install dependencies
pip install spacy
python -m spacy download pl_core_news_sm
Usage
from tokenizer import PolishToKenizer
# Load the tokenizer
tokenizer = PolishToKenizer.from_pretrained('vocab.json')
# Tokenize text
text = "Czytałem interesującą książkę o historii Polski."
# Get token strings
tokens = tokenizer.tokenize(text)
print(tokens) # ['▁Czyt', 'ał', 'em', '▁interesując', 'ą', '▁książkę', ...]
# Get token IDs
ids = tokenizer.tokens_to_ids(text)
print(ids) # [1234, 567, 89, ...]
# Decode back to text
decoded = tokenizer.ids_to_tokens(ids)
print(decoded) # "Czytałem interesującą książkę o historii Polski."
Morphological Awareness
Unlike standard BPE tokenizers, this implementation identifies morpheme boundaries before training:
| Word | Standard BPE | This Tokenizer | Linguistic Structure |
|---|---|---|---|
| Czytałem | Czyta|łem | Czyt|ał|em | stem + past + 1sg |
| napisałem | napis|ałem | na|pis|ał|em | prefix + stem + past + 1sg |
| książką | książ|ką | książk|ą | stem + instrumental |
Supported Morpheme Types
- Verb prefixes: na-, prze-, wy-, za-, pod-, etc. (20 prefixes)
- Verb endings: Infinitive, past tense, present tense, conditional, imperative, participles (108 morphemes)
- Noun endings: All 7 Polish cases in singular and plural (49 endings)
- Adjective endings: All cases, genders, comparative/superlative (44 morphemes)
- Adverb endings: -o, -e, -ie, -ej, -iej (5 endings)
Special Tokens
| Token | ID | Purpose |
|---|---|---|
<pad> |
0 | Padding |
<unk> |
1 | Unknown tokens |
<bos> |
2 | Beginning of sequence |
<eos> |
3 | End of sequence |
<mask> |
4 | Masked token (for MLM) |
Files
vocab.json- Trained vocabulary and tokenizer statetokenizer.py- Main tokenizer implementationtrie.py- Trie data structure for morpheme lookuppl_things/- Polish morphological resources (prefixes, endings, stop words)
Training Data
Trained on Polish Wikipedia (~300K unique words) with spaCy POS tagging for morpheme boundary detection.
Citation
If you use this tokenizer in your research, please cite:
@software{polish_morphological_tokenizer,
title = {Polish Morphological BPE Tokenizer},
year = {2026},
url = {https://huggingface.co/rafal-adamczyk/polish-morphological-tokenizer}
}
License
MIT License
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support