Instructions to use olaverse/diacnet-mini-2.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use olaverse/diacnet-mini-2.0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="olaverse/diacnet-mini-2.0")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSeq2SeqLM tokenizer = AutoTokenizer.from_pretrained("olaverse/diacnet-mini-2.0") model = AutoModelForSeq2SeqLM.from_pretrained("olaverse/diacnet-mini-2.0", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use olaverse/diacnet-mini-2.0 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "olaverse/diacnet-mini-2.0" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "olaverse/diacnet-mini-2.0", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/olaverse/diacnet-mini-2.0
- SGLang
How to use olaverse/diacnet-mini-2.0 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "olaverse/diacnet-mini-2.0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "olaverse/diacnet-mini-2.0", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "olaverse/diacnet-mini-2.0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "olaverse/diacnet-mini-2.0", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use olaverse/diacnet-mini-2.0 with Docker Model Runner:
docker model run hf.co/olaverse/diacnet-mini-2.0
diacnet-mini-2.0
diacnet-mini-2.0 Highlights
diacnet-mini-2.0 restores diacritics, tone marks and Arabic vowel marks in 11 languages from a single 300M byte-level model, with the following key features:
- Beats diactag-2.0 on 5 of 10 languages. With output alignment (below), its error rate is lower than diactag-2.0 on Polish, Portuguese, Spanish, French and Italian, and identical on Hausa.
- Yorùbá, rebuilt: error rate 0.174 → 0.081 against diacnet-1.1 (2.1× lower), and 9% of sentences exactly right (1.1: 0.8%).
- Now with Arabic: full diacritization (
<ara>, DER 0.110 on Classical Arabic) and a no-case-endings mode (<ara-nocase>, DER 0.082) from the same model. - Never changes your text: the alignment helper keeps every input letter and takes only the marks, so output compliance is 100% by construction. Raw output additionally fixes typos, if you want that.
- Language optional: with
<auto>, the error rate stays within 0.0012 of the explicit tag. - Meaning hints: add
[g: word=meaning]for words whose marks depend on meaning (Vietnamese ambiguous-word accuracy 0.53 → 0.78). - Partial input: marks already in the text are kept (99.9–100%) and the rest are filled in.
Model Overview
diacnet-mini-2.0 has the following features:
- Type: byte-level encoder-decoder (text to text), fine-tuned from google/byt5-small
- Architecture: ByT5-small: 12 encoder + 4 decoder layers, d_model 1472, 6 heads
- Parameters: 300M; weights 1.2 GB (fp32 safetensors)
- Languages and tags:
<yor>Yorùbá,<ibo>Igbo,<hau>Hausa,<vie>Vietnamese,<pol>Polish,<tur>Turkish,<por>Portuguese,<spa>Spanish,<fra>French,<ita>Italian,<ara>/<ara-nocase>Arabic,<auto> - Input:
<tag> [g: word=meaning] text, up to about 300 characters per call (the helper below splits longer text) - Output: the text with its diacritics restored
Quickstart
import torch
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
repo = "olaverse/diacnet-mini-2.0"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, dtype=torch.bfloat16).eval() # float32 on CPU
import difflib
import unicodedata as ud
_FOLD = str.maketrans("ɓɗƙƴđıłƁƊƘƳĐŁ", "bdkydilBDKYDL") # letters with no combining form
_LETTER = {"\u0653", "\u0654", "\u0655"} # Arabic madda / hamza are spelling, not marks
def _units(text):
units = [] # [letter, letter + marks]
for c in ud.normalize("NFD", text):
if units and ud.combining(c):
if c in _LETTER:
units[-1][0] += c
units[-1][1] += c
else:
units.append([c, c])
return [(ud.normalize("NFC", b).translate(_FOLD), ud.normalize("NFC", f)) for b, f in units]
def align(source, output):
"""`source` with the marks diacnet put on every letter it kept; letters it changed,
dropped or added fall back to `source`, so your text itself never changes."""
src, out = _units(source), _units(output)
res = [f for _, f in src]
sm = difflib.SequenceMatcher(None, [b for b, _ in src], [b for b, _ in out], autojunk=False)
for a, b, n in sm.get_matching_blocks():
res[a:a + n] = [f for _, f in out[b:b + n]]
return "".join(res)
def chunks(text, n=300):
out, cur = [], ""
for w in text.split(" "):
if cur and len(cur) + 1 + len(w) > n:
out.append(cur); cur = w
else:
cur = f"{cur} {w}" if cur else w
return out + [cur]
@torch.no_grad()
def diacritize(text, tag="auto", hint="", aligned=True):
parts = []
for c in chunks(text):
x = f"<{tag}> {hint + ' ' if hint else ''}{c}"
enc = tok(x, return_tensors="pt").to(model.device)
out = model.generate(**enc, max_new_tokens=2 * enc["input_ids"].shape[1] + 16, num_beams=1)
y = tok.decode(out[0], skip_special_tokens=True).strip()
parts.append(align(c, y) if aligned else y)
return " ".join(parts)
diacritize("O so fun ara re pe oun ko ni isoro kankan.", "yor")
# 'Ó sọ fún ara rẹ̀ pé òun kò ní ìṣòro kankan.'
diacritize("Dimkpa abuo a ga-ezute onwe ha ozo na mba Saudi Arabia.", "ibo")
# 'Dimkpa abụọ a ga-ezute onwe ha ọzọ na mba Saudi Arabia.'
diacritize("Cutar kan dauki tsawon kwana 14 zuwa 21 kafin ya warke.", "hau")
# 'Cutar kan ɗauki tsawon kwana 14 zuwa 21 kafin ya warke.'
diacritize("Karl bat dau lam bai tap ve nha cua minh.", "vie")
# 'Karl bắt đầu làm bài tập về nhà của mình.'
diacritize("Facebook zawiesil jedno z moich szesciu kont.", "pol")
# 'Facebook zawiesił jedno z moich sześciu kont.'
diacritize("No pude sujetarme a la cuerda mas tiempo.") # <auto>
# 'No pude sujetarme a la cuerda más tiempo.'
diacritize("وهذا قول مرغوب عنه .", "ara")
# 'وَهَذَا قَوْلٌ مَرْغُوبٌ عَنْهُ .'
For many sentences, batch inputs of similar length (padding="longest") and run on a GPU in bf16.
With the olaverse library
The olaverse library (v0.4.0+) wraps all of the above (chunking, batching, alignment, hints) in one call:
pip install "olaverse[deeplearning]"
from olaverse.nlp import Diacritizer
d = Diacritizer(model="diacnet-mini-2.0", lang="yor", device="cuda") # "cpu", "mps" or "auto"; bf16 on CUDA
d.restore("O so fun ara re pe oun ko ni isoro kankan.")
# 'Ó sọ fún ara rẹ̀ pé òun kò ní ìṣòro kankan.'
d.restore("وهذا قول مرغوب عنه .", lang="ara", case_endings=False)
# 'وَهَذَا قَوْل مَرْغُوب عَنْه .'
d.restore("Chi ay chi that su ranh vao nhung buoi toi sau khi da cho con ngu say.",
lang="vie", hints={"ranh": "free (time)"})
# 'Chị ấy chỉ thật sự rảnh vào những buổi tối sau khi đã cho con ngủ say.'
Diacritizer(model="diacnet-mini-2.0").restore("No pude sujetarme a la cuerda mas tiempo.") # no lang -> <auto>
# 'No pude sujetarme a la cuerda más tiempo.'
texts = ["O so fun ara re pe oun ko ni isoro kankan.", "وهذا قول مرغوب عنه .",
"No pude sujetarme a la cuerda mas tiempo."]
d.restore_batch(texts, lang=["yor", "ara", "auto"]) # one language per text, or one string for all
# ['Ó sọ fún ara rẹ̀ pé òun kò ní ìṣòro kankan.', 'وَهَذَا قَوْلٌ مَرْغُوبٌ عَنْهُ .',
# 'No pude sujetarme a la cuerda más tiempo.']
Output is aligned by default (aligned=False returns the raw output). hints also accepts a list of
"word=meaning" strings or a ready "[g: ...]" block.
Tags and Modes
Language tags and <auto>
Give the language tag when you know it. <auto> lets the model infer the language from the text; it costs at most
0.0012 DER on diacbench and nothing measurable on Yorùbá, Igbo, Hausa, Vietnamese or Arabic.
Arabic: <ara> and <ara-nocase>
<ara> restores everything, including the grammatical case ending on each word's last letter (for teaching,
classical and religious text, and text-to-speech). <ara-nocase> leaves word-final vowels off (shadda is kept),
which is how most modern Arabic is vowelled when it is vowelled at all.
diacritize("وهذا قول مرغوب عنه .", "ara-nocase") # 'وَهَذَا قَوْل مَرْغُوب عَنْه .'
Hamza letters (أ إ آ ؤ ئ) are spelling, not diacritics, and are kept as typed.
Meaning hints: [g: word=meaning]
Some words take different marks for different meanings (Vietnamese ranh: rảnh "free", rành "skilled", ranh
"border"; Yorùbá ogun: ogun "war", ogún "twenty"). A hint in English steers the choice; separate several hints
with |.
s = "Chi ay chi that su ranh vao nhung buoi toi sau khi da cho con ngu say."
diacritize(s, "vie", hint="[g: ranh=free (time)]")
# 'Chị ấy chỉ thật sự rảnh vào những buổi tối sau khi đã cho con ngủ say.'
diacritize(s, "vie")
# 'Chị ấy chỉ thật sự rành vào những buổi tối ...' (without the hint: "skilled")
Hints help most where the sentence alone is not enough (Vietnamese, Turkish, Polish). When a hint contradicts a clear sentence context, the model usually follows the context.
Partial input
Text that already carries some marks is completed, not undone: marks you typed are kept and the rest are filled in.
diacritize("Ki lo ṣe ti inu rẹ ko ni dun?", "yor") # underdots typed, tones missing
# 'Kí ló ṣe tí inú rẹ̀ kò ní dùn?'
Raw vs aligned output
aligned=True (recommended) guarantees that only marks change. aligned=False returns the model's own text,
which also corrects typos it was trained to fix (swapped, missing or doubled letters), but may change a letter
in a few percent of sentences.
Evaluation
Metrics. DER = share of mark-bearing letters (letters with more than one legal outcome in the language) whose marks differ from the reference. Strict scoring: an output that changes, adds or drops a letter counts every mark-bearing letter of that sentence as wrong. Exact = whole sentence correct. Compliance = output strips back to the input. Greedy decoding. Scored with the same code and letter definitions as the diactag-2.0 card, so the numbers compare directly.
diacbench (1,000 sentences per language)
Show table
DER on diacbench v1 references; diacnet outputs aligned.
| Language | diacnet-1.1 | diacnet-mini-2.0 | diacnet-2.0 | diactag-2.0 | diacnet-mini-2.0 on diacbench v2 | Exact (diacnet-mini-2.0) | Raw compliance (diacnet-mini-2.0) |
|---|---|---|---|---|---|---|---|
Yorùbá yor |
0.1743 | 0.0815 | 0.0698 | 0.0787 | 0.0761 | 0.087 | 0.993 |
Igbo ibo |
0.0219 | 0.0202 | 0.0185 | 0.0162 | 0.0177 | 0.419 | 0.937 |
Hausa hau |
0.0063 | 0.0044 | 0.0041 | 0.0044 | 0.0025 | 0.731 | 0.963 |
Vietnamese vie |
0.0316 | 0.0207 | 0.0143 | 0.0141 | 0.0204 | 0.523 | 1.000 |
Polish pol |
0.0055 | 0.0024 | 0.0014 | 0.0035 | 0.0023 | 0.954 | 0.998 |
Turkish tur |
0.0076 | 0.0042 | 0.0030 | 0.0035 | 0.0042 | 0.951 | 1.000 |
Portuguese por |
0.0042 | 0.0027 | 0.0021 | 0.0032 | 0.0026 | 0.929 | 1.000 |
Spanish spa |
0.0064 | 0.0031 | 0.0027 | 0.0038 | 0.0029 | 0.931 | 1.000 |
French fra |
0.0028 | 0.0013 | 0.0012 | 0.0017 | 0.0013 | 0.964 | 0.999 |
Italian ita |
0.0006 | 0.0003 | 0.0002 | 0.0005 | 0.0002 | 0.993 | 1.000 |
| Mean | 0.0261 | 0.0141 | 0.0117 | 0.0130 | 0.0130 | 0.748 |
Output alignment
Most of the raw error on Igbo, Hausa and Arabic comes from sentences where the model changed a letter, not from wrong marks. Alignment removes it.
Show table
| Test | Raw DER | Aligned DER | Raw compliance |
|---|---|---|---|
| Yorùbá | 0.0890 | 0.0815 | 0.993 |
| Igbo | 0.0907 | 0.0202 | 0.937 |
| Hausa | 0.0443 | 0.0044 | 0.963 |
| Vietnamese | 0.0207 | 0.0207 | 1.000 |
| Polish | 0.0049 | 0.0024 | 0.998 |
| Turkish | 0.0042 | 0.0042 | 1.000 |
| Portuguese | 0.0027 | 0.0027 | 1.000 |
| Spanish | 0.0031 | 0.0031 | 1.000 |
| French | 0.0021 | 0.0013 | 0.999 |
| Italian | 0.0003 | 0.0003 | 1.000 |
| Arabic, Classical (full) | 0.1114 | 0.1104 | 0.998 |
| Arabic, WikiNews (MSA) | 0.4267 | 0.2189 | 0.664 |
| Arabic, SadeedDiac-25 benchmark | 0.2949 | 0.1685 | 0.764 |
Arabic
Show table
DER over Arabic letters only, strict, on the same sentences for every system; ranked by WikiNews.
| System | WikiNews (MSA) | SadeedDiac-25 benchmark | Classical (diacbench ar) |
Classical (SadeedDiac-25 Fadel) |
|---|---|---|---|---|
| CATT-EO (Abjad AI) | 0.0396 | 0.0473 | 0.0140 | 0.0132 |
| diactag-2.0 | 0.0858 | 0.0994 | 0.0587 | 0.0625 |
| Shakkala v3 | 0.1056 | 0.1068 | 0.0365 | 0.0456 |
| Fine-Tashkeel | 0.1492 | 0.2165 | 0.0136 | 0.0207 |
| diacnet-2.0 (aligned) | 0.1542 | 0.1313 | 0.0955 | 0.0855 |
| diacnet-mini-2.0 (aligned) | 0.2189 | 0.1685 | 0.1104 | 0.1082 |
| Tashkeel-700M | 0.2961 | 0.3087 | 0.0392 | 0.0445 |
No-case-endings mode (<ara-nocase>) on Classical Arabic: DER 0.0821. diacnet was not trained on Tashkeela,
the corpus the Classical test sets come from (most other systems were), so the two Modern Standard Arabic columns
are the fairer comparison. Test sets: SadeedDiac-25
(WikiNewsTruth, 146 sentences; benchmark, 454; Fadel test, 600) and diacbench ar (1,000).
Meaning hints
Show table
Accuracy on the ambiguous word in 580 held-out sentences whose ambiguous spellings never appeared in training. The last column gives the hint of the word's other meaning instead and counts how often the output switches.
| Language | No hint | With hint | Switches with the other meaning's hint (n) |
|---|---|---|---|
| Vietnamese | 0.53 | 0.78 | 0.17 (92) |
| Turkish | 0.62 | 0.66 | 0.34 (32) |
| Polish | 0.89 | 0.98 | 0.00 (18) |
| Spanish | 1.00 | 1.00 | 0.00 (12) |
| French | 1.00 | 1.00 | – |
| Yorùbá | 0.41 | 0.41 | 0.36 (169) |
| Arabic | 0.72 | 0.72 | 0.43 (28) |
<auto> and partial input
Show tables
DER on the first 300 diacbench sentences per language (aligned):
| Language | Explicit tag | <auto> |
|---|---|---|
| Yorùbá | 0.0780 | 0.0781 |
| Igbo | 0.0171 | 0.0170 |
| Hausa | 0.0027 | 0.0025 |
| Vietnamese | 0.0213 | 0.0212 |
| Polish | 0.0022 | 0.0029 |
| Turkish | 0.0049 | 0.0049 |
| Portuguese | 0.0035 | 0.0039 |
| Spanish | 0.0033 | 0.0045 |
| French | 0.0017 | 0.0017 |
| Italian | 0.0004 | 0.0011 |
| Arabic | 0.1108 | 0.1110 |
Partial input, first 300 sentences (Yorùbá "underdots only": tones removed, underdots kept; "half the marks": a random half of the marked letters kept):
| Language | Marks given | DER (aligned) | DER from fully stripped input | Given marks kept |
|---|---|---|---|---|
| Hausa | half the marks | 0.0027 | 0.0027 | 1.000 |
| Igbo | half the marks | 0.0160 | 0.0171 | 1.000 |
| Vietnamese | half the marks | 0.0186 | 0.0213 | 1.000 |
| Yorùbá | half the marks | 0.0833 | 0.0780 | 0.999 |
| Yorùbá | underdots only | 0.0852 | 0.0780 | 1.000 |
Best Practices
- Use
align()unless you want typo correction: it costs nothing and removes letter changes. - Give the language tag when you know it;
<auto>is a safe default for mixed input. - Split long text into sentences or chunks of about 300 characters (the model was trained on units of that size).
- Greedy decoding (
num_beams=1),max_new_tokensabout twice the input length. - Choosing between Olaverse diacritizers:
| diacnet-2.0 | diacnet-mini-2.0 | diactag-2.0 | |
|---|---|---|---|
| Parameters | 582M | 300M | 37.9M |
| Approach | generative (text to text) | generative (text to text) | character tagger |
| Strongest on | Yorùbá and the European languages | close to diacnet-2.0 at half the size | Igbo, Vietnamese, Modern Standard Arabic |
| Extras | meaning hints, typo correction | meaning hints, typo correction | int8 ONNX for CPU, never alters letters |
Training
Trained from google/byt5-small for one epoch on about 2.66M (input, target) pairs: a
language-balanced 2.5M-pair sample of diacnet-1.1-train
(Yorùbá pairs kept only if well tone-marked, the cause of 1.1's Yorùbá regression) plus the teacher-marked set below,
repeated 4 times. AdamW, learning rate 3e-4, 3% warmup then linear decay, weight decay 0.01, effective batch 128
(16 × 8 gradient accumulation), bf16, inputs and targets up to 768 bytes, length-grouped batches. Inputs follow 1.1's format and
augmentations: <auto> 12%, partial marking 30%, meaning hints, character noise 7%.
Training Data and Licence
- diacnet-1.1-train: FineWeb-2 and Wikipedia sentences in the ten Latin-script languages (ODC-By 1.0 / CC BY-SA 4.0).
- About 37k teacher-marked pairs: real Yorùbá, Igbo, Hausa and Arabic news and web text re-marked in full by a large language model and kept only when its letters were unchanged, plus sentences built around words whose marks depend on meaning (Arabic in full and no-case-endings form). Classical Arabic from Fadel et al. (2019, MIT) was used only to confirm that rare spellings exist, not as training text.
- Benchmark sentences were removed from the training data.
Released under Apache-2.0.
Citation
@misc{diacnet-mini-2.0,
title = {diacnet-mini-2.0},
author = {Olaverse},
year = {2026},
url = {https://huggingface.co/olaverse/diacnet-mini-2.0}
}
- Downloads last month
- 442
Model tree for olaverse/diacnet-mini-2.0
Base model
google/byt5-small


