|
Download README.md from Mythologic/mark: direct link, hf CLI and curl.
- Browser
- Download file 12.6 kB
-
https://huggingface.co/Mythologic/mark/resolve/main/README.md
- Command line
-
hf download hf://Mythologic/mark/README.md
-
curl -L -o README.md https://huggingface.co/Mythologic/mark/resolve/main/README.md
12.6 kB
| language: | |
| - ak | |
| - ar | |
| - az | |
| - ca | |
| - cs | |
| - cy | |
| - ee | |
| - es | |
| - ff | |
| - fr | |
| - ga | |
| - gn | |
| - ha | |
| - he | |
| - hr | |
| - ht | |
| - hu | |
| - ig | |
| - ku | |
| - ln | |
| - lt | |
| - lv | |
| - mi | |
| - pl | |
| - pt | |
| - qu | |
| - ro | |
| - sk | |
| - sl | |
| - sm | |
| - sr | |
| - tk | |
| - tr | |
| - uz | |
| - vi | |
| - wo | |
| - yo | |
| license: apache-2.0 | |
| tags: | |
| - token-classification | |
| - diacritics | |
| - tone-restoration | |
| - vocalization | |
| - on-device | |
| - edge-ai | |
| - core-ml | |
| - onnx | |
| - webassembly | |
| - multilingual | |
| pipeline_tag: token-classification | |
| # Mark-38M | |
| A 38.7M-parameter byte-level Conv-Transformer that restores diacritics, tones, and vocalization marks across 37 languages without subword tokenizers or dictionary lookups: Akan (`ak`), Arabic (`ar`), Azerbaijani (`az`), Catalan (`ca`), Czech (`cs`), Welsh (`cy`), Ewe (`ee`), Spanish (`es`), Pulaar (`ff`), French (`fr`), Irish (`ga`), Guaraní (`gn`), Hausa (`ha`), Hebrew (`he`), Croatian (`hr`), Haitian Creole (`ht`), Hungarian (`hu`), Igbo (`ig`), Kurdish (`ku`), Lingala (`ln`), Lithuanian (`lt`), Latvian (`lv`), Māori (`mi`), Polish (`pl`), Portuguese (`pt`), Quechua (`qu`), Romanian (`ro`), Slovak (`sk`), Slovenian (`sl`), Samoan (`sm`), Serbian (`sr`), Turkmen (`tk`), Turkish (`tr`), Uzbek (`uz`), Vietnamese (`vi`), Wolof (`wo`), and Yorùbá (`yo`). | |
| On the official academic held-out **Yorùbá YAD test set** (3,330 sentences, 142k characters), it achieves a **15.88% Diacritic Error Rate (DER)**, **19.38% Word Error Rate (WER)**, and **5.58% Character Error Rate (CER)** with **0.0139% text corruption** (zero invented or dropped words), against 69.10% DER for the unmarked input. | |
| Across the 37-language joint evaluation suite, it achieves **93.69% macro marked-position accuracy** with a composite score of **0.8419**. Thirteen languages reach 100% Exact Match and 0.00% CER on benchmark probes. | |
| The model is exported into native on-device formats: a **41.77 MB INT8 ONNX graph** for CPU and WebAssembly, and a compiled **Core ML package** for the Apple Neural Engine. Quantized INT8 matches full-precision PyTorch with **99.64% character parity** (822 of 825 characters over the 74 evaluation probes). | |
| --- | |
| ## The Task | |
| Diacritic, tone, and vocalization restoration is not string rewriting or character translation. It is morphological disambiguation, clitic chain parsing, and syntactic agreement resolution operating under strict preservation constraints: | |
| 1. **Tone Languages** (Yorùbá, Akan, Ewe, Igbo, Lingala): Tones carry primary lexical and grammatical meaning. In Yorùbá, an unmarked syllable like `ba` can represent `bá` (to accompany/meet), `bà` (to perch/alight), or `ba` (to hide). Missing tones invert negation, tense, and aspect. | |
| 2. **Abjads** (Arabic, Modern Hebrew): Consonantal skeletons arrive without short vowels (*ḥarakāt*: *fatḥah*, *ḍammah*, *kasrah*, *sukūn*, *shaddah*) or case nunation (*tanwīn*). Syntax (*i'rab*), voice (active vs passive), and word semantics require contextual disambiguation across the sentence. | |
| 3. **Latin Orthographies** (Spanish, French, Portuguese, Turkish, Czech, Polish, Vietnamese, Romanian): Diacritics separate verb tenses (`comio` vs `comió`), nominal cases, and distinct phonemes (`c` vs `ç`, `s` vs `ş`, `a` vs `ă`). | |
| ### Why Bytes, Not Subwords or Characters | |
| Subword tokenizers (BPE, WordPiece) break down on diacritized text: | |
| * Unmarked input and marked output share base characters but differ in byte sequences. | |
| * Unicode normal forms (NFC precomposed vs NFD decomposed) split combining diacritics into orphan tokens, multiplying sequence lengths unpredictably. | |
| * Subword vocabularies cannot generalize to rare combining tone stacks (e.g. `M:0300+0301`). | |
| Mark operates directly on the **256 UTF-8 byte values**. Every input byte is assigned an operation tag from a discrete 1,073 tag vocabulary: | |
| * `KEEP`: Maintain the underlying byte. | |
| * `P:src->dst`: Substitute single character (e.g. `P:a->á`). | |
| * `M:hex`: Attach combining Unicode diacritic marks and normalize via canonical decomposition/composition (`NFC`). | |
| **Invariant Guarantee**: Every output character maps 1-to-1 to the underlying input stream. The model is structurally incapable of hallucinating, inserting, deleting, or reordering base words. | |
| --- | |
| ## Results | |
| ### Official Yorùbá YAD Test Benchmark (3,330 Sentences) | |
| Evaluated on the academic gold-standard [Yorùbá YAD test corpus](https://arxiv.org/abs/2004.14811) (Asahiah et al., Orife): | |
| | System | DER (Diacritic Error) | WER (Word Error) | CER (Char Error) | Hallucination / Invention | Inference Latency | | |
| | :--- | :---: | :---: | :---: | :---: | :---: | | |
| | **Mark-38M** | **15.88%** | **19.38%** | **5.58%** | **0.0139%** | **50.07 ms / sent** | | |
| | Unmarked Input (Identity Baseline) | 69.10% | 78.40% | 21.20% | 0.0000% | — | | |
| | Absolute Improvement | *-53.22%* | *-59.02%* | *-15.62%* | *Zero text drift* | *20 sent/sec (Apple Silicon)* | | |
| ### 37-Language Joint Accuracy Report | |
| Held-out evaluation report across all 37 languages: | |
| | Code | Language | Marked-Position Accuracy | Code | Language | Marked-Position Accuracy | | |
| | :--- | :--- | :---: | :--- | :--- | :---: | | |
| | `az` | Azerbaijani | **99.65%** | `ee` | Ewe | **97.39%** | | |
| | `fr` | French | **99.47%** | `sl` | Slovenian | **97.39%** | | |
| | `tr` | Turkish | **99.02%** | `sk` | Slovak | **97.21%** | | |
| | `gn` | Guaraní | **88.61%** | `lv` | Latvian | **96.86%** | | |
| | `es` | Spanish | **98.22%** | `ku` | Kurdish (Kurmanji) | **96.51%** | | |
| | `pt` | Portuguese | **98.09%** | `ig` | Igbo | **96.30%** | | |
| | `ga` | Irish | **98.07%** | `hu` | Hungarian | **95.60%** | | |
| | `sr` | Serbian | **98.07%** | `cs` | Czech | **95.47%** | | |
| | `pl` | Polish | **97.89%** | `ht` | Haitian Creole | **95.47%** | | |
| | `ak` | Akan (Twi) | **97.88%** | `ar` | Arabic | **94.86%** | | |
| | `hr` | Croatian | **97.84%** | `ca` | Catalan | **94.85%** | | |
| | `tk` | Turkmen | **97.80%** | `wo` | Wolof | **93.07%** | | |
| | `lt` | Lithuanian | **97.79%** | `ha` | Hausa | **92.18%** | | |
| | `ro` | Romanian | **97.41%** | `sm` | Samoan | **90.98%** | | |
| | `vi` | Vietnamese | **90.98%** | `qu` | Quechua | **89.78%** | | |
| | `he` | Hebrew | **89.40%** | `uz` | Uzbek | **89.15%** | | |
| | `yo` | Yorùbá | **87.39%** | `ff` | Pulaar (Fula) | **85.56%** | | |
| | `mi` | Māori | **81.15%** | `ln` | Lingala | **78.56%** | | |
| | `cy` | Welsh | **74.68%** | | | | | |
| * **Macro marked-position accuracy**: **93.69%** | |
| * **Worst-case language**: Welsh (`cy`) at **74.68%** | |
| * **Composite score**: **0.8419** | |
| --- | |
| ## Quantization & Edge Footprint | |
| | Format | Precision | File Size | Recommended Target | Latency (CPU / Apple NE) | | |
| | :--- | :--- | ---: | :--- | :---: | | |
| | `mark_int8.onnx` | Dynamic INT8 | **41.77 MB** | Edge CPU, Mobile, Browser (WASM) | 71.49 ms (4-thread CPU) | | |
| | `mark_fp16.onnx` | Float16 | **80.06 MB** | Mobile GPUs, WebGPU | 25.10 ms (GPU) | | |
| | `mark.mlpackage` | 8-bit Core ML | **42.10 MB** | Apple Neural Engine (iOS, macOS) | 12.40 ms (ANE) | | |
| | `mark_fp32.onnx` | Float32 | **158.84 MB** | Reference server baseline | 135.15 ms (1-thread CPU) | | |
| --- | |
| ## Install & Integration | |
| ### Python (ONNX Runtime) | |
| ```python | |
| from mark import restore | |
| # Standard inference | |
| restored = restore("Omode naa n kawe daadaa ni ile-iwe.", lang="yo") | |
| # -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé." | |
| # Calibrated anti-flicker mode for streaming keyboard input (margin >= 0.4) | |
| stable = restore("El nino comio jamon en la manana.", lang="es", margin=0.4) | |
| # -> "El niño comió jamón en la mañana." | |
| ``` | |
| ### iOS / macOS (Swift Package Manager) | |
| Add to your `Package.swift`: | |
| ```swift | |
| dependencies: [ | |
| .package(url: "https://huggingface.co/Mythologic/mark", branch: "main") | |
| ] | |
| ``` | |
| Declare dependency in your target: | |
| ```swift | |
| .product(name: "Mark", package: "mark") | |
| ``` | |
| Direct inference on Apple Neural Engine (Swift 6.3): | |
| ```swift | |
| import Mark | |
| let mark = try Mark() | |
| // Yorùbá | |
| let yoruba = try mark.restore("Omode naa n kawe daadaa ni ile-iwe.", lang: "yo") | |
| // -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé." | |
| // Arabic | |
| let arabic = try mark.restore("ذهب الولد الى المدرسة", lang: "ar") | |
| // -> "ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ" | |
| // Spanish (with calibrated margin: 0.4 for zero typing jitter) | |
| let spanish = try mark.restore("El nino comio jamon en la manana.", lang: "es", margin: 0.4) | |
| // -> "El niño comió jamón en la mañana." | |
| ``` | |
| ### Browser & Node.js (npm) | |
| Install via npm: | |
| ```bash | |
| npm i @mythologic/mark | |
| ``` | |
| Inference via ONNX Runtime Web (WASM / WebGPU): | |
| ```typescript | |
| import { Mark, restore } from "@mythologic/mark"; | |
| // Direct helper | |
| const yoruba = await restore("Omode naa n kawe daadaa ni ile-iwe.", "yo"); | |
| // -> Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé. | |
| const arabic = await restore("ذهب الولد الى المدرسة", "ar"); | |
| // -> ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ | |
| // Reusable instance with calibrated anti-flicker margin | |
| const mark = new Mark(); | |
| await mark.init(); | |
| const spanish = await mark.restore("El nino comio jamon en la manana.", "es", { margin: 0.4 }); | |
| // -> El niño comió jamón en la mañana. | |
| ``` | |
| --- | |
| ## Architecture | |
| | Attribute | Specification | | |
| | :--- | :--- | | |
| | **Parameters, total** | **38,735,217** (38.7M) | | |
| | **Encoder layers** | 16 Transformer encoder layers | | |
| | **Hidden dimension** | 384 | | |
| | **Intermediate dimension** | 1,536 (SwiGLU feed-forward) | | |
| | **Attention heads** | 8 heads (head dimension 48) | | |
| | **Positions** | Real-valued rotary embeddings (RoPE, $\theta = 10,000$, split-half NeoX form) | | |
| | **Convolution stem** | 3 parallel depthwise-separable 1D convolutions (kernels 3, 5, 7; dim 192) | | |
| | **Vocabulary** | 257 (256 raw UTF-8 bytes + language conditioning) | | |
| | **Classifier head** | Linear projection to 1,073 tag classes | | |
| | **Context window** | 512 bytes with sentence/whitespace sliding window | | |
| --- | |
| ## Training Setup | |
| * **Dataset volume**: ~786M tokens across 37 languages. | |
| * **Loss formulation**: Class-Balanced focal loss (`cb`, $w_{\text{KEEP}} = 0.25$, marked position weight 1.0). | |
| * **Sampling**: Temperature-scaled power law ($\tau = 0.5$) over language shards. | |
| * **Schedule**: 1,500 steps, batch size 1,024 sequences (512 bytes per sequence). | |
| * **Optimizer**: Muon (matrix trunk) + AdamW (embeddings & classifier). | |
| * **Hardware**: Dedicated high-throughput accelerator cluster. | |
| --- | |
| ## Limitations & Failure Modes | |
| 1. **Zero-Context Homographs**: Unmarked strings that represent multiple valid accented words in isolation (e.g. Spanish `si` [if] vs `sí` [yes], Arabic `qtr` -> `qaṭara` vs `qaṭr`) rely on local surrounding syntax within the 512-byte attention window. In total isolation, the model outputs the dominant corpus frequency. | |
| 2. **Tail Language Shard Volume**: High-resource shards (French, Spanish, Turkish, Arabic, Czech) exceed **95–99% accuracy**. Low-resource tail languages in the joint mixture (Welsh `cy` at 74.68%, Lingala `ln` at 78.56%) exhibit lower precision due to smaller corpus volume. | |
| 3. **Dialectal Writing**: Trained on standardized literary orthographies (Modern Standard Arabic, Standard Literary Yorùbá, Modern Hebrew). Informal dialectal orthography (e.g. Darija, Egyptian Arabic slang) exhibits degraded accuracy. | |
| 4. **Window Chunk Boundaries**: Sequences longer than 512 bytes are split at sentence and whitespace boundaries. Unpunctuated continuous text crossing a 512-byte boundary may experience slight boundary resolution loss. | |
| --- | |
| ## Files | |
| | File | Format | Size | Description | | |
| | :--- | :--- | ---: | :--- | | |
| | `mark_int8.onnx` | ONNX (INT8) | 41.77 MB | Dynamic INT8 quantized graph for CPU, Mobile, and WebAssembly | | |
| | `mark_fp16.onnx` | ONNX (FP16) | 80.06 MB | Half-precision graph for GPUs and Neural Engines | | |
| | `mark_fp32.onnx` | ONNX (FP32) | 158.84 MB | Full-precision reference model | | |
| | `mark.mlpackage.zip` | Core ML | 36.16 MB | Compiled Core ML package for Apple Neural Engine | | |
| | `config.json` | JSON | 1 KB | Model architectural hyperparameters | | |
| | `tags.json` | JSON | 40 KB | 1,073 tag operation mappings | | |
| | `languages.json` | JSON | 1 KB | 37 ISO 639-1 language codes | | |
| --- | |
| ## Author & Citation | |
| Developed by **Ainouche Abderahmane** and **Mythologic**. | |
| Released under the **Apache 2.0** License. | |
| ```bibtex | |
| @software{ainouche_mythologic_mark_2026, | |
| title = {Mark-38M: On-Device Multilingual Diacritic and Tone Restoration Across 37 Languages}, | |
| author = {Ainouche, Abderahmane and Mythologic}, | |
| year = {2026}, | |
| url = {https://huggingface.co/mythologic/mark}, | |
| note = {38.7M parameters; 93.69% macro marked-position accuracy across 37 languages} | |
| } | |
| ``` | |
| Part of [Mythologic](https://huggingface.co/mythologic). | |