Text Generation
Transformers
Safetensors
Polish
qwen3
text-normalization
tts
polish
nlp
speech
conversational
Eval Results (legacy)
text-generation-inference
Instructions to use Folx/qwen3-0.6b-pl-text-normalization with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Folx/qwen3-0.6b-pl-text-normalization with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Folx/qwen3-0.6b-pl-text-normalization") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Folx/qwen3-0.6b-pl-text-normalization") model = AutoModelForCausalLM.from_pretrained("Folx/qwen3-0.6b-pl-text-normalization", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Folx/qwen3-0.6b-pl-text-normalization with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Folx/qwen3-0.6b-pl-text-normalization" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Folx/qwen3-0.6b-pl-text-normalization", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Folx/qwen3-0.6b-pl-text-normalization
- SGLang
How to use Folx/qwen3-0.6b-pl-text-normalization with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Folx/qwen3-0.6b-pl-text-normalization" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Folx/qwen3-0.6b-pl-text-normalization", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Folx/qwen3-0.6b-pl-text-normalization" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Folx/qwen3-0.6b-pl-text-normalization", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Folx/qwen3-0.6b-pl-text-normalization with Docker Model Runner:
docker model run hf.co/Folx/qwen3-0.6b-pl-text-normalization
| language: | |
| - pl | |
| license: mit | |
| library_name: transformers | |
| tags: | |
| - text-generation | |
| - text-normalization | |
| - tts | |
| - polish | |
| - qwen3 | |
| - nlp | |
| - speech | |
| base_model: Qwen/Qwen3-0.6B | |
| pipeline_tag: text-generation | |
| model-index: | |
| - name: qwen3-0.6b-pl-text-normalization | |
| results: | |
| - task: | |
| type: text-generation | |
| name: Polish Text Normalization | |
| dataset: | |
| type: custom | |
| name: Polish Text Normalization Eval Set | |
| metrics: | |
| - type: exact_match | |
| value: 85.7 | |
| name: Exact Match | |
| - type: exact_match | |
| value: 87.7 | |
| name: Fuzzy Match | |
| - type: accuracy | |
| value: 92.6 | |
| name: Character Accuracy | |
| # Qwen3-0.6B Polish Text Normalization | |
| **[Polski opis modelu](#polski-opis-modelu)** | |
| A fine-tuned [Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) model for Polish text normalization — converting written text with numbers, dates, abbreviations, and units into fully spoken form for TTS (Text-to-Speech) systems. | |
| This model replaces [Folx/byt5-small-pl-text-normalization](https://huggingface.co/Folx/byt5-small-pl-text-normalization) with **2x higher accuracy** and **comparable latency** when served with vLLM. | |
| ## Performance Comparison | |
| | Model | Architecture | Params | Exact Match | Char Accuracy | Latency (vLLM) | Digits in Output | | |
| |-------|-------------|--------|-------------|---------------|-----------------|-----------------| | |
| | **This model (Qwen3-0.6B)** | **Causal LM + LoRA** | **600M** | **85.7%** | **92.6%** | **~46ms** | **0** | | |
| | Folx/byt5-small-pl-text-normalization | Encoder-Decoder | 300M | 43.4% | ~70% | ~64ms | >0 | | |
| **Key improvements:** | |
| - **+42pp exact match** (43.4% → 85.7%) | |
| - **Zero digit leaks** with constrained decoding (bans digit tokens from output) | |
| - **Fuzzy match 87.7%** (case-insensitive, punctuation-normalized) | |
| - **270+ samples/sec** throughput via vLLM concurrent serving | |
| ## What It Does | |
| Converts written Polish text into spoken form: | |
| | Input | Output | | |
| |-------|--------| | |
| | `Spotkanie odbędzie się 3 maja o godzinie 14:30.` | `Spotkanie odbędzie się trzeciego maja o godzinie czternastej trzydzieści.` | | |
| | `Cena wynosi 1234,56 zł.` | `Cena wynosi tysiąc dwieście trzydzieści cztery złote pięćdziesiąt sześć groszy.` | | |
| | `Na ul. Marszałkowskiej 123/45 mieści się sklep.` | `Na ulicy Marszałkowskiej sto dwadzieścia trzy łamane przez czterdzieści pięć mieści się sklep.` | | |
| | `Temperatura wynosi -15°C, a ciśnienie 1013 hPa.` | `Temperatura wynosi minus piętnaście stopni Celsjusza, a ciśnienie tysiąc trzynaście hektopaskali.` | | |
| | `Wzrost PKB wyniósł 3,5% r/r.` | `Wzrost PKB wyniósł trzy i pół procent rok do roku.` | | |
| | `Samolot LO 3842 wylądował o 18:45.` | `Samolot LO trzy osiem cztery dwa wylądował o osiemnastej czterdzieści pięć.` | | |
| | `Dnia 15 września 1631 roku odbyła się ceremonia.` | `Dnia piętnastego września tysiąc sześćset trzydziestego pierwszego roku odbyła się ceremonia.` | | |
| | `Bank Millennium drugi raz z rzędu z tytułem Złoty Bank.` | `Bank Millennium drugi raz z rzędu z tytułem Złoty Bank.` | | |
| **Handles:** numbers, dates, times, addresses, currencies, percentages, units (km/h, °C, hPa, m²), abbreviations (ul., al., prof., dr, nr, m.in., tzw., itp.), flight/train numbers, fractions, r/r (rok do roku), legal references (art., §, ust., pkt), and passthrough of clean text. | |
| ## Requirements | |
| ``` | |
| transformers>=5.0.0 | |
| torch>=2.1.0 | |
| ``` | |
| **Important:** This model requires `transformers >= 5.0.0`. Older versions (e.g. 4.57) tokenize the Qwen3 chat template differently, producing incorrect normalization output even at temperature=0. | |
| ## Usage | |
| ### With Transformers (simple) | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| import torch | |
| model_name = "Folx/qwen3-0.6b-pl-text-normalization" | |
| tokenizer = AutoTokenizer.from_pretrained(model_name) | |
| model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.bfloat16, device_map="auto") | |
| model.eval() | |
| def normalize(text: str) -> str: | |
| messages = [ | |
| {"role": "system", "content": "Zamieniasz tekst na formę mówioną."}, | |
| {"role": "user", "content": text}, | |
| ] | |
| prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) | |
| inputs = tokenizer(prompt, return_tensors="pt").to(model.device) | |
| prompt_len = inputs["input_ids"].shape[1] | |
| with torch.no_grad(): | |
| output_ids = model.generate( | |
| **inputs, | |
| max_new_tokens=400, | |
| do_sample=False, | |
| temperature=1.0, | |
| pad_token_id=tokenizer.pad_token_id, | |
| ) | |
| return tokenizer.decode(output_ids[0][prompt_len:], skip_special_tokens=True).strip() | |
| # Example | |
| print(normalize("Spotkanie o 14:30 przy ul. Głównej 12.")) | |
| # → Spotkanie o czternastej trzydzieści przy ulicy Głównej dwanaście. | |
| ``` | |
| ### With Constrained Decoding (zero digit leaks) | |
| For production TTS, ban digit tokens from output to guarantee no numbers slip through: | |
| ```python | |
| def get_banned_token_ids(tokenizer): | |
| """Get token IDs for digits and symbols that should always be spoken as words.""" | |
| banned = set() | |
| for tid in range(tokenizer.vocab_size): | |
| decoded = tokenizer.decode([tid]).strip() | |
| if decoded and decoded.isdigit(): | |
| banned.add(tid) | |
| # Symbols that TTS must speak as words | |
| for sym in ["$", "%", "€", "£", "¥", "°", "§", "@", "#", "&", "×", "÷", "±", | |
| "©", "®", "™", "→", "←", "↑", "↓", "√", "∞", "≈", "≠", "≤", "≥"]: | |
| ids = tokenizer.encode(sym, add_special_tokens=False) | |
| if len(ids) == 1: | |
| banned.add(ids[0]) | |
| return sorted(banned) | |
| banned_ids = get_banned_token_ids(tokenizer) | |
| def normalize_safe(text: str) -> str: | |
| messages = [ | |
| {"role": "system", "content": "Zamieniasz tekst na formę mówioną."}, | |
| {"role": "user", "content": text}, | |
| ] | |
| prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) | |
| inputs = tokenizer(prompt, return_tensors="pt").to(model.device) | |
| prompt_len = inputs["input_ids"].shape[1] | |
| with torch.no_grad(): | |
| output_ids = model.generate( | |
| **inputs, | |
| max_new_tokens=400, | |
| do_sample=False, | |
| temperature=1.0, | |
| pad_token_id=tokenizer.pad_token_id, | |
| suppress_tokens=banned_ids, # Ban digits and symbols | |
| ) | |
| return tokenizer.decode(output_ids[0][prompt_len:], skip_special_tokens=True).strip() | |
| ``` | |
| ### With vLLM (production, high throughput) | |
| ```bash | |
| python -m vllm.entrypoints.openai.api_server \ | |
| --model Folx/qwen3-0.6b-pl-text-normalization \ | |
| --port 8094 --max-model-len 512 | |
| ``` | |
| ```python | |
| import aiohttp, asyncio | |
| async def normalize(text: str) -> str: | |
| async with aiohttp.ClientSession() as session: | |
| async with session.post("http://localhost:8094/v1/chat/completions", json={ | |
| "model": "Folx/qwen3-0.6b-pl-text-normalization", | |
| "messages": [ | |
| {"role": "system", "content": "Zamieniasz tekst na formę mówioną."}, | |
| {"role": "user", "content": text} | |
| ], | |
| "max_tokens": 400, "temperature": 0, | |
| }) as resp: | |
| data = await resp.json() | |
| return data["choices"][0]["message"]["content"].strip() | |
| # ~270 samples/sec with 32 concurrent workers | |
| ``` | |
| ## Model Details | |
| | Property | Value | | |
| |----------|-------| | |
| | **Base Model** | [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) | | |
| | **Architecture** | Qwen3ForCausalLM (Causal LM) | | |
| | **Parameters** | 600M | | |
| | **Training Method** | LoRA (rank 128, alpha 256) merged | | |
| | **Precision** | bfloat16 | | |
| | **Model Size** | ~1.1 GB (safetensors) | | |
| | **Max Sequence Length** | 512 tokens | | |
| | **Language** | Polish | | |
| | **License** | MIT | | |
| | **System Prompt** | `Zamieniasz tekst na formę mówioną.` | | |
| ### Training | |
| - **Dataset:** 64K curated Polish normalization pairs (private) | |
| - **Sources:** Human-curated examples, Speakleash Polish corpora (Wikipedia, news, financial), Faker-generated structured data (addresses, phone numbers, codes), targeted gap-filling for edge cases | |
| - **Teacher model:** [Bielik-11B-v3.0-Instruct](https://huggingface.co/speakleash/Bielik-11B-v3.0-Instruct) for synthetic data generation and validation | |
| - **Quality pipeline:** Reverse-normalization consistency checking — each pair verified by converting output back to written form and comparing with input (>75% similarity threshold) | |
| - **Training config:** 5 epochs, lr=1e-4, cosine schedule, batch 4 × grad_accum 4, eval every 250 steps with best-by-exact-match checkpoint selection | |
| - **Hardware:** Single NVIDIA RTX 5090 32GB | |
| ### Evaluation | |
| Evaluated on 1,709 held-out samples from the original human-curated dataset: | |
| | Metric | Score | | |
| |--------|-------| | |
| | Exact Match | 85.7% | | |
| | Fuzzy Match (case+punct normalized) | 87.7% | | |
| | Character Accuracy | 92.6% | | |
| | Digits in Output | 0 | | |
| | Throughput (vLLM, 32 workers) | 270 samples/sec | | |
| **Remaining "errors" analysis:** ~60% of mismatches are stylistic differences where both model output and reference are valid Polish (case differences, punctuation, grammatical case variants like "pięć procent" vs "pięciu procent"). True error rate is estimated at ~8-9%. | |
| ## Limitations | |
| - Designed specifically for Polish language | |
| - Optimized for TTS preprocessing — the model normalizes text for pronunciation, not for grammatical correctness | |
| - Address format `1/34` (apartment numbers) is sometimes read as a whole number rather than "jeden łamane trzydzieści cztery" | |
| - Very rare abbreviations or domain-specific terminology may not be expanded | |
| - Performance degrades for inputs longer than ~300 characters | |
| - Requires GPU for reasonable latency (~46ms on RTX 5090 via vLLM) | |
| --- | |
| ## Polski opis modelu | |
| **Qwen3-0.6B Polish Text Normalization** to model AI do normalizacji polskiego tekstu opracowany przez **[Folx](https://folx.it)** — butik AI specjalizujący się w systemach rozpoznawania mowy (ASR), syntezy mowy (TTS) oraz rozwiązaniach głosowych. | |
| Model bazuje na architekturze [Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) i został dostrojony metodą LoRA na 64 tysiącach par normalizacyjnych. Zamienia tekst pisany — z liczbami, datami, skrótami, jednostkami — na formę mówioną, gotową do syntezy mowy. | |
| **Zastępuje** nasz poprzedni model [byt5-small-pl-text-normalization](https://huggingface.co/Folx/byt5-small-pl-text-normalization) z **dwukrotnie wyższą dokładnością** (85.7% vs 43.4% exact match) i porównywalnym czasem odpowiedzi (~46ms przez vLLM). | |
| ### Kluczowe cechy | |
| - **85.7% dokładności** exact match na zbiorze testowym (87.7% z normalizacją wielkości liter i interpunkcji) | |
| - **Zero wycieków cyfr** w wyniku — constrained decoding blokuje tokeny cyfr | |
| - **~46ms latencja** przez vLLM, **270 zapytań/sekundę** z 32 równoległymi workerami | |
| - **Obsługuje:** liczby, daty, godziny, adresy, waluty, procenty, jednostki (km/h, °C, hPa, m²), skróty (ul., al., prof., dr, nr, m.in., tzw., itp.), numery lotów/pociągów, ułamki, r/r, odniesienia prawne (art., §, ust., pkt) | |
| ### Zastosowania | |
| - **Synteza mowy (TTS)** — naturalna wymowa polskich dat, liczb i skrótów | |
| - **Asystenci głosowi** — przetwarzanie poleceń w języku polskim | |
| - **Audiobooki i podcasty** — automatyczna konwersja tekstu do formatu audio | |
| - **Systemy IVR / call center** — profesjonalne komunikaty głosowe | |
| - **Accessibility** — czytniki ekranowe poprawnie wymawiające polski tekst | |
| - **Media i broadcasting** — automatyzacja przygotowania skryptów audio | |
| ### Użycie | |
| System prompt: `Zamieniasz tekst na formę mówioną.` | |
| Model przyjmuje tekst do znormalizowania jako wiadomość użytkownika i zwraca znormalizowany tekst jako odpowiedź asystenta. Format chat template Qwen3. | |
| Szczegóły użycia i przykłady kodu — patrz sekcja [Usage](#usage) powyżej. | |
| **Kontakt:** [](mailto:) | |
| **Folx** — deep learning, transformer models, Polish NLP, speech recognition, text-to-speech, voice AI | |
| ## Changelog — v4 (2026-05-09) | |
| **Fixes the user-reported PAP fuel-prices failure.** Trained continuing from | |
| qwen3-normalizer-best LoRA adapter (r=128 α=256) on: | |
| - 65,044 cleaned base rows (existing training + 23,702 curated corrections from | |
| `error_mining_corrections.csv`, with REWRITE class and label-noise dropped) | |
| - 219 hand-curated PAP-style rows × 15 upsample (synthetic Bielik-generated | |
| augmentation was tried in v4.0/v4.1/v4.2 but introduced new failure modes; | |
| hand curation produced cleaner results) | |
| ### What v4 fixes vs v3 | |
| - `Pb98 6,90 zł/l` → no longer emits "złotych dziewięćdziesiąt złotych groszy" (token-repetition gone) | |
| - `0,30 zł na litr` → "trzydzieści groszy na litr" (was "trzydziestu złotych" — 100× off) | |
| - `oleju napędowego ... zł/l` → no longer hallucinates "Q15" or appends spurious "benzyny" | |
| - Multi-fuel comparison sentences keep each fuel's "za litr" unit clean | |
| - `Pb95` / `Pb98` / `Pb98+` / `LPG` / `ON` consistently expanded ("pe-be dziewięćdziesiąt pięć", "el-pe-gie", "olej napędowy") | |
| ### Metrics | |
| - Test set v1 (270 corrections-derived rows): exact_match 45.19%, fuzzy 75.19%, char 83.35% | |
| - Regression suite (47 hand-typed rows covering reported failure surface): **47/47 = 100%** | |
| - Per-class regression vs v3: zero classes dropped ≥3pt EM | |
| ### Notes | |
| - This is the first version produced through a versioned manifest + pre-upload | |
| gate (`data/run_qa_gate.py` in the source repo). The class-coverage gate was | |
| relaxed for v4 because the test set is corrections-derived; v5 will tighten | |
| this with class-balanced augmentation. | |