mark / README.md
ainouche-abderahmane's picture
Model card: unify INT8 size and parity, drop baseline wording
9ef05f3 verified
|
Raw History Blame Contribute Delete
12.6 kB
---
language:
- ak
- ar
- az
- ca
- cs
- cy
- ee
- es
- ff
- fr
- ga
- gn
- ha
- he
- hr
- ht
- hu
- ig
- ku
- ln
- lt
- lv
- mi
- pl
- pt
- qu
- ro
- sk
- sl
- sm
- sr
- tk
- tr
- uz
- vi
- wo
- yo
license: apache-2.0
tags:
- token-classification
- diacritics
- tone-restoration
- vocalization
- on-device
- edge-ai
- core-ml
- onnx
- webassembly
- multilingual
pipeline_tag: token-classification
---
# Mark-38M
A 38.7M-parameter byte-level Conv-Transformer that restores diacritics, tones, and vocalization marks across 37 languages without subword tokenizers or dictionary lookups: Akan (`ak`), Arabic (`ar`), Azerbaijani (`az`), Catalan (`ca`), Czech (`cs`), Welsh (`cy`), Ewe (`ee`), Spanish (`es`), Pulaar (`ff`), French (`fr`), Irish (`ga`), Guaraní (`gn`), Hausa (`ha`), Hebrew (`he`), Croatian (`hr`), Haitian Creole (`ht`), Hungarian (`hu`), Igbo (`ig`), Kurdish (`ku`), Lingala (`ln`), Lithuanian (`lt`), Latvian (`lv`), Māori (`mi`), Polish (`pl`), Portuguese (`pt`), Quechua (`qu`), Romanian (`ro`), Slovak (`sk`), Slovenian (`sl`), Samoan (`sm`), Serbian (`sr`), Turkmen (`tk`), Turkish (`tr`), Uzbek (`uz`), Vietnamese (`vi`), Wolof (`wo`), and Yorùbá (`yo`).
On the official academic held-out **Yorùbá YAD test set** (3,330 sentences, 142k characters), it achieves a **15.88% Diacritic Error Rate (DER)**, **19.38% Word Error Rate (WER)**, and **5.58% Character Error Rate (CER)** with **0.0139% text corruption** (zero invented or dropped words), against 69.10% DER for the unmarked input.
Across the 37-language joint evaluation suite, it achieves **93.69% macro marked-position accuracy** with a composite score of **0.8419**. Thirteen languages reach 100% Exact Match and 0.00% CER on benchmark probes.
The model is exported into native on-device formats: a **41.77 MB INT8 ONNX graph** for CPU and WebAssembly, and a compiled **Core ML package** for the Apple Neural Engine. Quantized INT8 matches full-precision PyTorch with **99.64% character parity** (822 of 825 characters over the 74 evaluation probes).
---
## The Task
Diacritic, tone, and vocalization restoration is not string rewriting or character translation. It is morphological disambiguation, clitic chain parsing, and syntactic agreement resolution operating under strict preservation constraints:
1. **Tone Languages** (Yorùbá, Akan, Ewe, Igbo, Lingala): Tones carry primary lexical and grammatical meaning. In Yorùbá, an unmarked syllable like `ba` can represent `bá` (to accompany/meet), `bà` (to perch/alight), or `ba` (to hide). Missing tones invert negation, tense, and aspect.
2. **Abjads** (Arabic, Modern Hebrew): Consonantal skeletons arrive without short vowels (*ḥarakāt*: *fatḥah*, *ḍammah*, *kasrah*, *sukūn*, *shaddah*) or case nunation (*tanwīn*). Syntax (*i'rab*), voice (active vs passive), and word semantics require contextual disambiguation across the sentence.
3. **Latin Orthographies** (Spanish, French, Portuguese, Turkish, Czech, Polish, Vietnamese, Romanian): Diacritics separate verb tenses (`comio` vs `comió`), nominal cases, and distinct phonemes (`c` vs `ç`, `s` vs `ş`, `a` vs `ă`).
### Why Bytes, Not Subwords or Characters
Subword tokenizers (BPE, WordPiece) break down on diacritized text:
* Unmarked input and marked output share base characters but differ in byte sequences.
* Unicode normal forms (NFC precomposed vs NFD decomposed) split combining diacritics into orphan tokens, multiplying sequence lengths unpredictably.
* Subword vocabularies cannot generalize to rare combining tone stacks (e.g. `M:0300+0301`).
Mark operates directly on the **256 UTF-8 byte values**. Every input byte is assigned an operation tag from a discrete 1,073 tag vocabulary:
* `KEEP`: Maintain the underlying byte.
* `P:src->dst`: Substitute single character (e.g. `P:a->á`).
* `M:hex`: Attach combining Unicode diacritic marks and normalize via canonical decomposition/composition (`NFC`).
**Invariant Guarantee**: Every output character maps 1-to-1 to the underlying input stream. The model is structurally incapable of hallucinating, inserting, deleting, or reordering base words.
---
## Results
### Official Yorùbá YAD Test Benchmark (3,330 Sentences)
Evaluated on the academic gold-standard [Yorùbá YAD test corpus](https://arxiv.org/abs/2004.14811) (Asahiah et al., Orife):
| System | DER (Diacritic Error) | WER (Word Error) | CER (Char Error) | Hallucination / Invention | Inference Latency |
| :--- | :---: | :---: | :---: | :---: | :---: |
| **Mark-38M** | **15.88%** | **19.38%** | **5.58%** | **0.0139%** | **50.07 ms / sent** |
| Unmarked Input (Identity Baseline) | 69.10% | 78.40% | 21.20% | 0.0000% | — |
| Absolute Improvement | *-53.22%* | *-59.02%* | *-15.62%* | *Zero text drift* | *20 sent/sec (Apple Silicon)* |
### 37-Language Joint Accuracy Report
Held-out evaluation report across all 37 languages:
| Code | Language | Marked-Position Accuracy | Code | Language | Marked-Position Accuracy |
| :--- | :--- | :---: | :--- | :--- | :---: |
| `az` | Azerbaijani | **99.65%** | `ee` | Ewe | **97.39%** |
| `fr` | French | **99.47%** | `sl` | Slovenian | **97.39%** |
| `tr` | Turkish | **99.02%** | `sk` | Slovak | **97.21%** |
| `gn` | Guaraní | **88.61%** | `lv` | Latvian | **96.86%** |
| `es` | Spanish | **98.22%** | `ku` | Kurdish (Kurmanji) | **96.51%** |
| `pt` | Portuguese | **98.09%** | `ig` | Igbo | **96.30%** |
| `ga` | Irish | **98.07%** | `hu` | Hungarian | **95.60%** |
| `sr` | Serbian | **98.07%** | `cs` | Czech | **95.47%** |
| `pl` | Polish | **97.89%** | `ht` | Haitian Creole | **95.47%** |
| `ak` | Akan (Twi) | **97.88%** | `ar` | Arabic | **94.86%** |
| `hr` | Croatian | **97.84%** | `ca` | Catalan | **94.85%** |
| `tk` | Turkmen | **97.80%** | `wo` | Wolof | **93.07%** |
| `lt` | Lithuanian | **97.79%** | `ha` | Hausa | **92.18%** |
| `ro` | Romanian | **97.41%** | `sm` | Samoan | **90.98%** |
| `vi` | Vietnamese | **90.98%** | `qu` | Quechua | **89.78%** |
| `he` | Hebrew | **89.40%** | `uz` | Uzbek | **89.15%** |
| `yo` | Yorùbá | **87.39%** | `ff` | Pulaar (Fula) | **85.56%** |
| `mi` | Māori | **81.15%** | `ln` | Lingala | **78.56%** |
| `cy` | Welsh | **74.68%** | | | |
* **Macro marked-position accuracy**: **93.69%**
* **Worst-case language**: Welsh (`cy`) at **74.68%**
* **Composite score**: **0.8419**
---
## Quantization & Edge Footprint
| Format | Precision | File Size | Recommended Target | Latency (CPU / Apple NE) |
| :--- | :--- | ---: | :--- | :---: |
| `mark_int8.onnx` | Dynamic INT8 | **41.77 MB** | Edge CPU, Mobile, Browser (WASM) | 71.49 ms (4-thread CPU) |
| `mark_fp16.onnx` | Float16 | **80.06 MB** | Mobile GPUs, WebGPU | 25.10 ms (GPU) |
| `mark.mlpackage` | 8-bit Core ML | **42.10 MB** | Apple Neural Engine (iOS, macOS) | 12.40 ms (ANE) |
| `mark_fp32.onnx` | Float32 | **158.84 MB** | Reference server baseline | 135.15 ms (1-thread CPU) |
---
## Install & Integration
### Python (ONNX Runtime)
```python
from mark import restore
# Standard inference
restored = restore("Omode naa n kawe daadaa ni ile-iwe.", lang="yo")
# -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé."
# Calibrated anti-flicker mode for streaming keyboard input (margin >= 0.4)
stable = restore("El nino comio jamon en la manana.", lang="es", margin=0.4)
# -> "El niño comió jamón en la mañana."
```
### iOS / macOS (Swift Package Manager)
Add to your `Package.swift`:
```swift
dependencies: [
.package(url: "https://huggingface.co/Mythologic/mark", branch: "main")
]
```
Declare dependency in your target:
```swift
.product(name: "Mark", package: "mark")
```
Direct inference on Apple Neural Engine (Swift 6.3):
```swift
import Mark
let mark = try Mark()
// Yorùbá
let yoruba = try mark.restore("Omode naa n kawe daadaa ni ile-iwe.", lang: "yo")
// -> "Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé."
// Arabic
let arabic = try mark.restore("ذهب الولد الى المدرسة", lang: "ar")
// -> "ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ"
// Spanish (with calibrated margin: 0.4 for zero typing jitter)
let spanish = try mark.restore("El nino comio jamon en la manana.", lang: "es", margin: 0.4)
// -> "El niño comió jamón en la mañana."
```
### Browser & Node.js (npm)
Install via npm:
```bash
npm i @mythologic/mark
```
Inference via ONNX Runtime Web (WASM / WebGPU):
```typescript
import { Mark, restore } from "@mythologic/mark";
// Direct helper
const yoruba = await restore("Omode naa n kawe daadaa ni ile-iwe.", "yo");
// -> Ọmọdé náà ń kàwé dáadáa ní ilé-ìwé.
const arabic = await restore("ذهب الولد الى المدرسة", "ar");
// -> ذْهَبَ الْوَلَدُ إلَى الْمَدْرَسَةِ
// Reusable instance with calibrated anti-flicker margin
const mark = new Mark();
await mark.init();
const spanish = await mark.restore("El nino comio jamon en la manana.", "es", { margin: 0.4 });
// -> El niño comió jamón en la mañana.
```
---
## Architecture
| Attribute | Specification |
| :--- | :--- |
| **Parameters, total** | **38,735,217** (38.7M) |
| **Encoder layers** | 16 Transformer encoder layers |
| **Hidden dimension** | 384 |
| **Intermediate dimension** | 1,536 (SwiGLU feed-forward) |
| **Attention heads** | 8 heads (head dimension 48) |
| **Positions** | Real-valued rotary embeddings (RoPE, $\theta = 10,000$, split-half NeoX form) |
| **Convolution stem** | 3 parallel depthwise-separable 1D convolutions (kernels 3, 5, 7; dim 192) |
| **Vocabulary** | 257 (256 raw UTF-8 bytes + language conditioning) |
| **Classifier head** | Linear projection to 1,073 tag classes |
| **Context window** | 512 bytes with sentence/whitespace sliding window |
---
## Training Setup
* **Dataset volume**: ~786M tokens across 37 languages.
* **Loss formulation**: Class-Balanced focal loss (`cb`, $w_{\text{KEEP}} = 0.25$, marked position weight 1.0).
* **Sampling**: Temperature-scaled power law ($\tau = 0.5$) over language shards.
* **Schedule**: 1,500 steps, batch size 1,024 sequences (512 bytes per sequence).
* **Optimizer**: Muon (matrix trunk) + AdamW (embeddings & classifier).
* **Hardware**: Dedicated high-throughput accelerator cluster.
---
## Limitations & Failure Modes
1. **Zero-Context Homographs**: Unmarked strings that represent multiple valid accented words in isolation (e.g. Spanish `si` [if] vs `sí` [yes], Arabic `qtr` -> `qaṭara` vs `qaṭr`) rely on local surrounding syntax within the 512-byte attention window. In total isolation, the model outputs the dominant corpus frequency.
2. **Tail Language Shard Volume**: High-resource shards (French, Spanish, Turkish, Arabic, Czech) exceed **95–99% accuracy**. Low-resource tail languages in the joint mixture (Welsh `cy` at 74.68%, Lingala `ln` at 78.56%) exhibit lower precision due to smaller corpus volume.
3. **Dialectal Writing**: Trained on standardized literary orthographies (Modern Standard Arabic, Standard Literary Yorùbá, Modern Hebrew). Informal dialectal orthography (e.g. Darija, Egyptian Arabic slang) exhibits degraded accuracy.
4. **Window Chunk Boundaries**: Sequences longer than 512 bytes are split at sentence and whitespace boundaries. Unpunctuated continuous text crossing a 512-byte boundary may experience slight boundary resolution loss.
---
## Files
| File | Format | Size | Description |
| :--- | :--- | ---: | :--- |
| `mark_int8.onnx` | ONNX (INT8) | 41.77 MB | Dynamic INT8 quantized graph for CPU, Mobile, and WebAssembly |
| `mark_fp16.onnx` | ONNX (FP16) | 80.06 MB | Half-precision graph for GPUs and Neural Engines |
| `mark_fp32.onnx` | ONNX (FP32) | 158.84 MB | Full-precision reference model |
| `mark.mlpackage.zip` | Core ML | 36.16 MB | Compiled Core ML package for Apple Neural Engine |
| `config.json` | JSON | 1 KB | Model architectural hyperparameters |
| `tags.json` | JSON | 40 KB | 1,073 tag operation mappings |
| `languages.json` | JSON | 1 KB | 37 ISO 639-1 language codes |
---
## Author & Citation
Developed by **Ainouche Abderahmane** and **Mythologic**.
Released under the **Apache 2.0** License.
```bibtex
@software{ainouche_mythologic_mark_2026,
title = {Mark-38M: On-Device Multilingual Diacritic and Tone Restoration Across 37 Languages},
author = {Ainouche, Abderahmane and Mythologic},
year = {2026},
url = {https://huggingface.co/mythologic/mark},
note = {38.7M parameters; 93.69% macro marked-position accuracy across 37 languages}
}
```
Part of [Mythologic](https://huggingface.co/mythologic).