File size: 8,998 Bytes
d6db43f
 
 
 
 
 
 
 
 
 
e7a8143
 
 
 
 
d6db43f
 
 
e7a8143
d6db43f
 
 
 
 
 
e7a8143
d6db43f
e7a8143
d6db43f
e7a8143
 
d6db43f
e7a8143
 
 
 
 
 
 
 
 
d6db43f
 
e7a8143
d6db43f
e7a8143
d6db43f
e7a8143
 
 
 
 
d6db43f
e7a8143
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8cedb24
e7a8143
 
 
 
 
 
 
 
 
 
 
 
 
d6db43f
 
 
e7a8143
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
317c6f8
e7a8143
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7942060
 
 
 
 
 
 
6542816
7942060
 
 
317c6f8
7942060
 
 
 
 
 
 
 
 
e7a8143
 
 
 
 
d6db43f
 
e7a8143
 
 
 
 
 
 
 
 
1646460
 
 
 
d6db43f
 
 
e7a8143
 
 
d6db43f
e7a8143
 
 
 
 
 
 
 
 
 
 
d6db43f
 
 
e7a8143
d6db43f
e7a8143
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
---
language:
- la
license: mit
tags:
- latin
- token-classification
- inverse-text-normalization
- casing
- punctuation
- capitalization
- punctuation-restoration
- asr-post-processing
- classical-latin
pipeline_tag: token-classification
datasets:
- njand/latin-asr-post-processing-dataset
model-index:
- name: Latin ASR Post-Processor
  results:
  - task:
      type: token-classification
      name: Inverse Text Normalization
    dataset:
      type: njand/latin-asr-post-processing-dataset
      name: Latin ASR Post-Processing Dataset
    metrics:
    - name: Macro F1
      type: f1
      value: 0.6523
    - name: Overall Accuracy
      type: accuracy
      value: 0.9216
    - name: Precision
      type: precision
      value: 0.6228
    - name: Recall
      type: recall
      value: 0.6943
base_model:
- latincy/latin-bert
---

# Latin ASR Post-Processor (Casing & Punctuation Restoration)

An Inverse Text Normalization (ITN) transformer model fine-tuned to convert unformatted, raw Latin Automatic Speech Recognition (ASR) outputs into fully formatted, classical Latin text. It simultaneously restores capitalization and trailing punctuation using a **14-class composite sequence labeling schema**.

---

### πŸ“Œ Quick Links
- **Live Demo:** [Gradio Interface](https://huggingface.co/spaces/njand/latin-asr-demo)
- **Source Code:** [GitHub Repository](https://github.com/njand/latin-asr-postprocess)
- **Base Model:** [`latincy/latin-bert`](https://huggingface.co/latincy/latin-bert)
- **Dataset:** [`njand/latin-asr-post-processing-dataset`](https://huggingface.co/datasets/njand/latin-asr-post-processing-dataset)

---

## πŸ› οΈ Pipeline Architecture & Preprocessing

This model is intended to be used directly downstream of the acoustic model [`njand/wav2vec2-xls-r-latin`](https://huggingface.co/njand/wav2vec2-xls-r-latin). 

Because raw ASR models emit stream-of-consciousness text (lowercased, space-separated, and unpunctuated), the text must pass through an input normalization pipeline before being fed into this model for casing and punctuation restoration.

```text
+-----------------------+     +-------------------------------+     +-------------------------------+
|  Raw Audio Waveform   | --> |  njand/wav2vec2-xls-r-latin   | --> | Preprocessing & Normalization |
+-----------------------+     +-------------------------------+     +-------------------------------+
                                                                                    |
                                                                                    v
+-----------------------+     +-------------------------------+     +-------------------------------+
| Formatted Text Output | <-- |   Latin ASR Post-Processor    | <-- |  Custom CLTK Tokenization     |
+-----------------------+     +-------------------------------+     +-------------------------------+
```

### Input Preprocessing Requirements

To prepare raw transcript outputs for inference, apply the following sequence of transformations:

1. **Macron Stripping:** Remove all vowel length diacritics (e.g., *ā, Δ“, Δ«, ō, Ε«, Θ³* β†’ *a, e, i, o, u, y*).
2. **Orthographic Standardization:** Standardize consonant/vowel variants:
   * Convert *j* β†’ *i* and *v* β†’ *u*.
   * Handle orthographic exceptions (e.g., *ejicio* β†’ *eicio*).
3. **Custom CLTK Word Tokenization:** Run the normalized string through a version of the **[CLTK (Classical Language Toolkit v0)](https://github.com/cltk/cltk/tree/v0/cltk/tokenize/latin)** Latin word tokenizer.
   > *Note:* Because official CLTK v0 tokenization scripts are unmaintained, a bespoke implementation of the tokenizer was executed dynamically during training preprocessing rather than being pre-applied to the static dataset.

---

## πŸš€ Quickstart & Inference Utility

Below is a complete Python script demonstrating how to prepare raw ASR output and run inference using the post-processing pipeline.

```python
from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline

PUNCT_MAP = {
    "NONE": "",
    "COMMA": ",",
    "PERIOD": ".",
    "SEMICOLON": ";",
    "COLON": ":",
    "QUESTION": "?",
    "EXCLAMATION": "!",
}

def format_token(word: str, tag: str) -> str:
    """Applies composite ITN tag (e.g., 'TITLE_COMMA') to a word token."""
    parts = tag.split("_")
    if len(parts) != 2:
        return word

    casing, punct = parts[0], parts[1]

    if casing == "TITLE":
        word = word.capitalize()
    elif casing == "LOWER":
        word = word.lower()

    return f"{word}{PUNCT_MAP.get(punct, '')}"

def restore_latin_text(pipe, raw_text: str) -> str:
    """Runs inference and reconstructs formatted Latin text."""
    predictions = pipe(raw_text, aggregation_strategy="first")
    formatted_words = [
        format_token(pred["word"].strip(" "), pred["entity_group"])
        for pred in predictions
    ]
    return " ".join(formatted_words)

# 1. Load pipeline
model_id = "njand/latin-asr-postprocessor"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForTokenClassification.from_pretrained(model_id)

itn_pipe = pipeline("token-classification", model=model, tokenizer=tokenizer)

# 2. Test reconstruction with preprocessed ASR output
raw_asr_input = "gallia est omnis divisa in partes tres quarum unam incolunt belgae"
print(restore_latin_text(itn_pipe, raw_asr_input))
# Output: "Gallia est omnis divisa in partes tres, quarum unam incolunt Belgae."

```

---

## 🏷️ Composite Label Schema

Target labels utilize a **14-class composite sequence schema** that pairs Casing state with Trailing Punctuation state:

$$\text{Label} = \text{Casing} \times \text{Punctuation}$$

* **Casing Tags (2):** `LOWER`, `TITLE`
* **Punctuation Tags (7):** `NONE`, `COMMA`, `PERIOD`, `SEMICOLON`, `COLON`, `QUESTION`, `EXCLAMATION`

---

## βš™οΈ Model Variants & Optimization

To facilitate production deployment on CPU-based infrastructure, this repository provides the model in three formats:

| Format | Precision | File Size | Latency (P50) | Recommended Use Case |
| :--- | :--- | :--- | :--- | :--- |
| **PyTorch** | FP32 | 443 MB | - | Training, fine-tuning, and PyTorch pipelines |
| **ONNX** | FP32 | 443 MB | 21.5 ms | **Production (Maximum Accuracy)** |
| **ONNX Quantized** | INT8 | **188 MB** | **9.7 ms** | **Low-Latency & Edge CPU** |

> **Performance vs. Precision Trade-off:**
> While INT8 dynamic quantization yields a **2.2Γ— speedup (P50)** and cuts RAM usage by **57.5%**, top-line accuracy (92.16% β†’ 91.03%) hides a severe drop in macro performance:
> * **Macro F1 Collapse:** Drops from **65.19% to 50.00%**. Dynamic weight quantization compresses logit decision boundaries for rare token tags.
> * **Punctuation Degradation:** Punctuation F1 falls **10.66 percentage points** (69.09% β†’ 58.43%), causing increased missing or misclassified commas, colons, and sentence boundaries.
> * **Casing Stability:** Capitalization F1 remains mostly intact (91.71% β†’ 89.38%).
>
> **Recommendation:** Use **ONNX FP32** for production pipelines where text formatting and punctuation precision are critical. Use **ONNX INT8** in latency-critical environments where speed and memory constraints outweigh exact punctuation recovery.
> 

---

## πŸ“Š Benchmarks & Performance (Epoch 7 - Best Checkpoint)

Evaluated on a 95/5 train/holdout split across diverse Classical Latin literary and historical corpora.

### Overall Summary Metrics

| Metric | Score |
| --- | --- |
| **Overall Accuracy** | **92.16%** |
| **Macro F1** | **0.6523** |
| **Precision** | **62.28%** |
| **Recall** | **69.43%** |
| **Validation Loss** | **0.2592** |

### 🎯 Sub-Task Breakdown

| Task | Accuracy | F1 Score |
| --- | --- | --- |
| **Casing Restoration** | **98.35%** | **0.9171** |
| **Punctuation Insertion** | **93.67%** | **0.6908** |

---

## πŸ“ˆ Training Progression

The model was trained over 9 epochs fine-tuning [`latincy/latin-bert`](https://www.google.com/url?sa=E&source=gmail&q=https://huggingface.co/latincy/latin-bert). Model weights from **Epoch 7** were selected based on optimal overall F1.

| Epoch | Train Loss | Val Loss | Overall F1 | Overall Acc | Casing Acc | Punct Acc |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | 0.6284 | 0.2883 | 0.6097 | 91.05% | 98.08% | 92.79% |
| 2 | 0.5617 | 0.2701 | 0.6275 | 91.60% | 98.19% | 93.26% |
| 3 | 0.5150 | 0.2618 | 0.6388 | 91.91% | 98.27% | 93.50% |
| 4 | 0.4855 | 0.2607 | 0.6455 | 92.04% | 98.30% | 93.60% |
| 5 | 0.4652 | 0.2579 | 0.6440 | 92.12% | 98.33% | 93.65% |
| 6 | 0.4502 | 0.2582 | 0.6462 | 92.15% | 98.34% | 93.66% |
| **7** | **0.4351** | **0.2592** | **0.6523** | **92.16%** | **98.35%** | **93.67%** |
| 8 | 0.4262 | 0.2598 | 0.6495 | 92.24% | 98.36% | 93.75% |
| 9 | 0.4172 | 0.2601 | 0.6507 | 92.21% | 98.37% | 93.71% |

---

## ⚑ Hardware & Environmental Footprint

* **Hardware Infrastructure:** NVIDIA L4 GPU via Modal
* **Training Time:** 3.82 hours
* **Estimated Carbon Emissions:** 0.1797 kg COβ‚‚eq