Polish Morphological BPE Tokenizer

A morphologically-aware BPE (Byte-Pair Encoding) tokenizer for Polish that respects morpheme boundaries during tokenization.

Key Features

  • Morpheme-aware: Prevents merges across morphological boundaries (prefixes, suffixes, case endings)
  • SentencePiece-style: Uses ▁ (U+2581) to mark word boundaries for perfect reconstruction
  • POS-aware: Uses spaCy for part-of-speech tagging to apply correct morpheme rules per word type
  • Vocabulary: 30,000 tokens trained on Polish Wikipedia

Installation

# Clone the repository
git clone https://huggingface.co/rafal-adamczyk/polish-morphological-tokenizer
cd polish-morphological-tokenizer

# Install dependencies
pip install spacy
python -m spacy download pl_core_news_sm

Usage

from tokenizer import PolishToKenizer

# Load the tokenizer
tokenizer = PolishToKenizer.from_pretrained('vocab.json')

# Tokenize text
text = "Czytałem interesującą książkę o historii Polski."

# Get token strings
tokens = tokenizer.tokenize(text)
print(tokens)  # ['▁Czyt', 'ał', 'em', '▁interesując', 'ą', '▁książkę', ...]

# Get token IDs
ids = tokenizer.tokens_to_ids(text)
print(ids)  # [1234, 567, 89, ...]

# Decode back to text
decoded = tokenizer.ids_to_tokens(ids)
print(decoded)  # "Czytałem interesującą książkę o historii Polski."

Morphological Awareness

Unlike standard BPE tokenizers, this implementation identifies morpheme boundaries before training:

Word Standard BPE This Tokenizer Linguistic Structure
Czytałem Czyta|łem Czyt|ał|em stem + past + 1sg
napisałem napis|ałem na|pis|ał|em prefix + stem + past + 1sg
książką książ|ką książk|ą stem + instrumental

Supported Morpheme Types

  • Verb prefixes: na-, prze-, wy-, za-, pod-, etc. (20 prefixes)
  • Verb endings: Infinitive, past tense, present tense, conditional, imperative, participles (108 morphemes)
  • Noun endings: All 7 Polish cases in singular and plural (49 endings)
  • Adjective endings: All cases, genders, comparative/superlative (44 morphemes)
  • Adverb endings: -o, -e, -ie, -ej, -iej (5 endings)

Special Tokens

Token ID Purpose
<pad> 0 Padding
<unk> 1 Unknown tokens
<bos> 2 Beginning of sequence
<eos> 3 End of sequence
<mask> 4 Masked token (for MLM)

Files

  • vocab.json - Trained vocabulary and tokenizer state
  • tokenizer.py - Main tokenizer implementation
  • trie.py - Trie data structure for morpheme lookup
  • pl_things/ - Polish morphological resources (prefixes, endings, stop words)

Training Data

Trained on Polish Wikipedia (~300K unique words) with spaCy POS tagging for morpheme boundary detection.

Citation

If you use this tokenizer in your research, please cite:

@software{polish_morphological_tokenizer,
  title = {Polish Morphological BPE Tokenizer},
  year = {2026},
  url = {https://huggingface.co/rafal-adamczyk/polish-morphological-tokenizer}
}

License

MIT License

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support