--- language: - ak - ar - az - ca - cs - cy - ee - es - ff - fr - ga - gn - ha - he - hr - ht - hu - ig - ku - ln - lt - lv - mi - pl - pt - qu - ro - sk - sl - sm - sr - tk - tr - uz - vi - wo - yo license: apache-2.0 tags: - token-classification - diacritics - tone-restoration - vocalization - on-device - edge-ai - core-ml - onnx - webassembly - multilingual pipeline_tag: token-classification --- # Mark-38M A 38.7M-parameter byte-level Conv-Transformer that restores diacritics, tones, and vocalization marks across 37 languages without subword tokenizers or dictionary lookups: Akan (`ak`), Arabic (`ar`), Azerbaijani (`az`), Catalan (`ca`), Czech (`cs`), Welsh (`cy`), Ewe (`ee`), Spanish (`es`), Pulaar (`ff`), French (`fr`), Irish (`ga`), Guaraní (`gn`), Hausa (`ha`), Hebrew (`he`), Croatian (`hr`), Haitian Creole (`ht`), Hungarian (`hu`), Igbo (`ig`), Kurdish (`ku`), Lingala (`ln`), Lithuanian (`lt`), Latvian (`lv`), Māori (`mi`), Polish (`pl`), Portuguese (`pt`), Quechua (`qu`), Romanian (`ro`), Slovak (`sk`), Slovenian (`sl`), Samoan (`sm`), Serbian (`sr`), Turkmen (`tk`), Turkish (`tr`), Uzbek (`uz`), Vietnamese (`vi`), Wolof (`wo`), and Yorùbá (`yo`). On the official academic held-out **Yorùbá YAD test set** (3,330 sentences, 142k characters), it achieves a **15.88% Diacritic Error Rate (DER)**, **19.38% Word Error Rate (WER)**, and **5.58% Character Error Rate (CER)** with **0.0139% text corruption** (zero invented or dropped words), against 69.10% DER for the unmarked input. Across the 37-language joint evaluation suite, it achieves **93.69% macro marked-position accuracy** with a composite score of **0.8419**. Thirteen languages reach 100% Exact Match and 0.00% CER on benchmark probes. The model is exported into native on-device formats: a **41.77 MB INT8 ONNX graph** for CPU and WebAssembly, and a compiled **Core ML package** for the Apple Neural Engine. Quantized INT8 matches full-precision PyTorch with **99.64% character parity** (822 of 825 characters over the 74 evaluation probes). --- ## The Task Diacritic, tone, and vocalization restoration is not string rewriting or character translation. It is morphological disambiguation, clitic chain parsing, and syntactic agreement resolution operating under strict preservation constraints: 1. **Tone Languages** (Yorùbá, Akan, Ewe, Igbo, Lingala): Tones carry primary lexical and grammatical meaning. In Yorùbá, an unmarked syllable like `ba` can represent `bá` (to accompany/meet), `bà` (to perch/alight), or `ba` (to hide). Missing tones invert negation, tense, and aspect. 2. **Abjads** (Arabic, Modern Hebrew): Consonantal skeletons arrive without short vowels (*ḥarakāt*: *fatḥah*, *ḍammah*, *kasrah*, *sukūn*, *shaddah*) or case nunation (*tanwīn*). Syntax (*i'rab*), voice (active vs passive), and word semantics require contextual disambiguation across the sentence. 3. **Latin Orthographies** (Spanish, French, Portuguese, Turkish, Czech, Polish, Vietnamese, Romanian): Diacritics separate verb tenses (`comio` vs `comió`), nominal cases, and distinct phonemes (`c` vs `ç`, `s` vs `ş`, `a` vs `ă`). ### Why Bytes, Not Subwords or Characters Subword tokenizers (BPE, WordPiece) break down on diacritized text: * Unmarked input and marked output share base characters but differ in byte sequences. * Unicode normal forms (NFC precomposed vs NFD decomposed) split combining diacritics into orphan tokens, multiplying sequence lengths unpredictably. * Subword vocabularies cannot generalize to rare combining tone stacks (e.g. `M:0300+0301`). Mark operates directly on the **256 UTF-8 byte values**. Every input byte is assigned an operation tag from a discrete 1,073 tag vocabulary: * `KEEP`: Maintain the underlying byte. * `P:src->dst`: Substitute single character (e.g. `P:a->á`). * `M:hex`: Attach combining Unicode diacritic marks and normalize via canonical decomposition/composition (`NFC`). **Invariant Guarantee**: Every output character maps 1-to-1 to the underlying input stream. The model is structurally incapable of hallucinating, inserting, deleting, or reordering base words. --- ## Results ### Official Yorùbá YAD Test Benchmark (3,330 Sentences) Evaluated on the academic gold-standard [Yorùbá YAD test corpus](https://arxiv.org/abs/2004.14811) (Asahiah et al., Orife): | System | DER (Diacritic Error) | WER (Word Error) | CER (Char Error) | Hallucination / Invention | Inference Latency | | :--- | :---: | :---: | :---: | :---: | :---: | | **Mark-38M** | **15.88%** | **19.38%** | **5.58%** | **0.0139%** | **50.07 ms / sent** | | Unmarked Input (Identity Baseline) | 69.10% | 78.40% | 21.20% | 0.0000% | — | | Absolute Improvement | *-53.22%* | *-59.02%* | *-15.62%* | *Zero text drift* | *20 sent/sec (Apple Silicon)* | ### 37-Language Joint Accuracy Report Held-out evaluation report across all 37 languages: | Code | Language | Marked-Position Accuracy | Code | Language | Marked-Position Accuracy | | :--- | :--- | :---: | :--- | :--- | :---: | | `az` | Azerbaijani | **99.65%** | `ee` | Ewe | **97.39%** | | `fr` | French | **99.47%** | `sl` | Slovenian | **97.39%** | | `tr` | Turkish | **99.02%** | `sk` | Slovak | **97.21%** | | `gn` | Guaraní | **88.61%** | `lv` | Latvian | **96.86%** | | `es` | Spanish | **98.22%** | `ku` | Kurdish (Kurmanji) | **96.51%** | | `pt` | Portuguese | **98.09%** | `ig` | Igbo | **96.30%** | | `ga` | Irish | **98.07%** | `hu` | Hungarian | **95.60%** | | `sr` | Serbian | **98.07%** | `cs` | Czech | **95.47%** | | `pl` | Polish | **97.89%** | `ht` | Haitian Creole | **95.47%** | | `ak` | Akan (Twi) | **97.88%** | `ar` | Arabic | **94.86%** | | `hr` | Croatian | **97.84%** | `ca` | Catalan | **94.85%** | | `tk` | Turkmen | **97.80%** | `wo` | Wolof | **93.07%** | | `lt` | Lithuanian | **97.79%** | `ha` | Hausa | **92.18%** | | `ro` | Romanian | **97.41%** | `sm` | Samoan | **90.98%** | | `vi` | Vietnamese | **90.98%** | `qu` | Quechua | **89.78%** | | `he` | Hebrew | **89.40%** | `uz` | Uzbek | **89.15%** | | `yo` | Yorùbá | **87.39%** | `ff` | Pulaar (Fula) | **85.56%** | | `mi` | Māori | **81.15%** | `ln` | Lingala | **78.56%** | | `cy` | Welsh | **74.68%** | | | | * **Macro marked-position accuracy**: **93.69%** * **Worst-case language**: Welsh (`cy`) at **74.68%** * **Composite score**: **0.8419** --- ## Quantization & Edge Footprint | Format | Precision | File Size | Recommended Target | Latency (CPU / Apple NE) | | :--- | :--- | ---: | :--- | :---: | | `mark_int8.onnx` | Dynamic INT8 | **41.77 MB** | Edge CPU, Mobile, Browser (WASM) | 71.49 ms (4-thread CPU) | | `mark_fp16.onnx` | Float16 | **80.06 MB** | Mobile GPUs, WebGPU | 25.10 ms (GPU) | | `mark.mlpackage` | 8-bit Core ML | **42.10 MB** | Apple Neural Engine (iOS, macOS) | 12.40 ms (ANE) | | `mark_fp32.onnx` | Float32 | **158.84 MB** | Reference server baseline | 135.15 ms (1-thread CPU) | --- ## Install & Integration ### Python (ONNX Runtime) ```python from mark import restore # Standard inference restored = restore("Omode naa n kawe daadaa ni ile-iwe.", lang="yo") # -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé." # Calibrated anti-flicker mode for streaming keyboard input (margin >= 0.4) stable = restore("El nino comio jamon en la manana.", lang="es", margin=0.4) # -> "El niño comió jamón en la mañana." ``` ### iOS / macOS (Swift Package Manager) Add to your `Package.swift`: ```swift dependencies: [ .package(url: "https://huggingface.co/Mythologic/mark", branch: "main") ] ``` Declare dependency in your target: ```swift .product(name: "Mark", package: "mark") ``` Direct inference on Apple Neural Engine (Swift 6.3): ```swift import Mark let mark = try Mark() // Yorùbá let yoruba = try mark.restore("Omode naa n kawe daadaa ni ile-iwe.", lang: "yo") // -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé." // Arabic let arabic = try mark.restore("ذهب الولد الى المدرسة", lang: "ar") // -> "ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ" // Spanish (with calibrated margin: 0.4 for zero typing jitter) let spanish = try mark.restore("El nino comio jamon en la manana.", lang: "es", margin: 0.4) // -> "El niño comió jamón en la mañana." ``` ### Browser & Node.js (npm) Install via npm: ```bash npm i @mythologic/mark ``` Inference via ONNX Runtime Web (WASM / WebGPU): ```typescript import { Mark, restore } from "@mythologic/mark"; // Direct helper const yoruba = await restore("Omode naa n kawe daadaa ni ile-iwe.", "yo"); // -> Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé. const arabic = await restore("ذهب الولد الى المدرسة", "ar"); // -> ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ // Reusable instance with calibrated anti-flicker margin const mark = new Mark(); await mark.init(); const spanish = await mark.restore("El nino comio jamon en la manana.", "es", { margin: 0.4 }); // -> El niño comió jamón en la mañana. ``` --- ## Architecture | Attribute | Specification | | :--- | :--- | | **Parameters, total** | **38,735,217** (38.7M) | | **Encoder layers** | 16 Transformer encoder layers | | **Hidden dimension** | 384 | | **Intermediate dimension** | 1,536 (SwiGLU feed-forward) | | **Attention heads** | 8 heads (head dimension 48) | | **Positions** | Real-valued rotary embeddings (RoPE, $\theta = 10,000$, split-half NeoX form) | | **Convolution stem** | 3 parallel depthwise-separable 1D convolutions (kernels 3, 5, 7; dim 192) | | **Vocabulary** | 257 (256 raw UTF-8 bytes + language conditioning) | | **Classifier head** | Linear projection to 1,073 tag classes | | **Context window** | 512 bytes with sentence/whitespace sliding window | --- ## Training Setup * **Dataset volume**: ~786M tokens across 37 languages. * **Loss formulation**: Class-Balanced focal loss (`cb`, $w_{\text{KEEP}} = 0.25$, marked position weight 1.0). * **Sampling**: Temperature-scaled power law ($\tau = 0.5$) over language shards. * **Schedule**: 1,500 steps, batch size 1,024 sequences (512 bytes per sequence). * **Optimizer**: Muon (matrix trunk) + AdamW (embeddings & classifier). * **Hardware**: Dedicated high-throughput accelerator cluster. --- ## Limitations & Failure Modes 1. **Zero-Context Homographs**: Unmarked strings that represent multiple valid accented words in isolation (e.g. Spanish `si` [if] vs `sí` [yes], Arabic `qtr` -> `qaṭara` vs `qaṭr`) rely on local surrounding syntax within the 512-byte attention window. In total isolation, the model outputs the dominant corpus frequency. 2. **Tail Language Shard Volume**: High-resource shards (French, Spanish, Turkish, Arabic, Czech) exceed **95–99% accuracy**. Low-resource tail languages in the joint mixture (Welsh `cy` at 74.68%, Lingala `ln` at 78.56%) exhibit lower precision due to smaller corpus volume. 3. **Dialectal Writing**: Trained on standardized literary orthographies (Modern Standard Arabic, Standard Literary Yorùbá, Modern Hebrew). Informal dialectal orthography (e.g. Darija, Egyptian Arabic slang) exhibits degraded accuracy. 4. **Window Chunk Boundaries**: Sequences longer than 512 bytes are split at sentence and whitespace boundaries. Unpunctuated continuous text crossing a 512-byte boundary may experience slight boundary resolution loss. --- ## Files | File | Format | Size | Description | | :--- | :--- | ---: | :--- | | `mark_int8.onnx` | ONNX (INT8) | 41.77 MB | Dynamic INT8 quantized graph for CPU, Mobile, and WebAssembly | | `mark_fp16.onnx` | ONNX (FP16) | 80.06 MB | Half-precision graph for GPUs and Neural Engines | | `mark_fp32.onnx` | ONNX (FP32) | 158.84 MB | Full-precision reference model | | `mark.mlpackage.zip` | Core ML | 36.16 MB | Compiled Core ML package for Apple Neural Engine | | `config.json` | JSON | 1 KB | Model architectural hyperparameters | | `tags.json` | JSON | 40 KB | 1,073 tag operation mappings | | `languages.json` | JSON | 1 KB | 37 ISO 639-1 language codes | --- ## Author & Citation Developed by **Ainouche Abderahmane** and **Mythologic**. Released under the **Apache 2.0** License. ```bibtex @software{ainouche_mythologic_mark_2026, title = {Mark-38M: On-Device Multilingual Diacritic and Tone Restoration Across 37 Languages}, author = {Ainouche, Abderahmane and Mythologic}, year = {2026}, url = {https://huggingface.co/mythologic/mark}, note = {38.7M parameters; 93.69% macro marked-position accuracy across 37 languages} } ``` Part of [Mythologic](https://huggingface.co/mythologic).