Latin ASR Post-Processor (Casing & Punctuation Restoration)

An Inverse Text Normalization (ITN) transformer model fine-tuned to convert unformatted, raw Latin Automatic Speech Recognition (ASR) outputs into fully formatted, classical Latin text. It simultaneously restores capitalization and trailing punctuation using a 14-class composite sequence labeling schema.


📌 Quick Links


🛠️ Pipeline Architecture & Preprocessing

This model is intended to be used directly downstream of the acoustic model njand/wav2vec2-xls-r-latin.

Because raw ASR models emit stream-of-consciousness text (lowercased, space-separated, and unpunctuated), the text must pass through an input normalization pipeline before being fed into this model for casing and punctuation restoration.

+-----------------------+     +-------------------------------+     +-------------------------------+
|  Raw Audio Waveform   | --> |  njand/wav2vec2-xls-r-latin   | --> | Preprocessing & Normalization |
+-----------------------+     +-------------------------------+     +-------------------------------+
                                                                                    |
                                                                                    v
+-----------------------+     +-------------------------------+     +-------------------------------+
| Formatted Text Output | <-- |   Latin ASR Post-Processor    | <-- |  Custom CLTK Tokenization     |
+-----------------------+     +-------------------------------+     +-------------------------------+

Input Preprocessing Requirements

To prepare raw transcript outputs for inference, apply the following sequence of transformations:

  1. Macron Stripping: Remove all vowel length diacritics (e.g., ā, ē, ī, ō, ū, ȳa, e, i, o, u, y).
  2. Orthographic Standardization: Standardize consonant/vowel variants:
    • Convert ji and vu.
    • Handle orthographic exceptions (e.g., ejicioeicio).
  3. Custom CLTK Word Tokenization: Run the normalized string through a version of the CLTK (Classical Language Toolkit v0) Latin word tokenizer.

    Note: Because official CLTK v0 tokenization scripts are unmaintained, a bespoke implementation of the tokenizer was executed dynamically during training preprocessing rather than being pre-applied to the static dataset.


🚀 Quickstart & Inference Utility

Below is a complete Python script demonstrating how to prepare raw ASR output and run inference using the post-processing pipeline.

from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline

PUNCT_MAP = {
    "NONE": "",
    "COMMA": ",",
    "PERIOD": ".",
    "SEMICOLON": ";",
    "COLON": ":",
    "QUESTION": "?",
    "EXCLAMATION": "!",
}

def format_token(word: str, tag: str) -> str:
    """Applies composite ITN tag (e.g., 'TITLE_COMMA') to a word token."""
    parts = tag.split("_")
    if len(parts) != 2:
        return word

    casing, punct = parts[0], parts[1]

    if casing == "TITLE":
        word = word.capitalize()
    elif casing == "LOWER":
        word = word.lower()

    return f"{word}{PUNCT_MAP.get(punct, '')}"

def restore_latin_text(pipe, raw_text: str) -> str:
    """Runs inference and reconstructs formatted Latin text."""
    predictions = pipe(raw_text, aggregation_strategy="first")
    formatted_words = [
        format_token(pred["word"].strip(" "), pred["entity_group"])
        for pred in predictions
    ]
    return " ".join(formatted_words)

# 1. Load pipeline
model_id = "njand/latin-asr-postprocessor"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForTokenClassification.from_pretrained(model_id)

itn_pipe = pipeline("token-classification", model=model, tokenizer=tokenizer)

# 2. Test reconstruction with preprocessed ASR output
raw_asr_input = "gallia est omnis divisa in partes tres quarum unam incolunt belgae"
print(restore_latin_text(itn_pipe, raw_asr_input))
# Output: "Gallia est omnis divisa in partes tres, quarum unam incolunt Belgae."

🏷️ Composite Label Schema

Target labels utilize a 14-class composite sequence schema that pairs Casing state with Trailing Punctuation state:

Label=Casing×Punctuation\text{Label} = \text{Casing} \times \text{Punctuation}

  • Casing Tags (2): LOWER, TITLE
  • Punctuation Tags (7): NONE, COMMA, PERIOD, SEMICOLON, COLON, QUESTION, EXCLAMATION

⚙️ Model Variants & Optimization

To facilitate production deployment on CPU-based infrastructure, this repository provides the model in three formats:

Format Precision File Size Latency (P50) Recommended Use Case
PyTorch FP32 443 MB - Training, fine-tuning, and PyTorch pipelines
ONNX FP32 443 MB 21.5 ms Production (Maximum Accuracy)
ONNX Quantized INT8 188 MB 9.7 ms Low-Latency & Edge CPU

Performance vs. Precision Trade-off: While INT8 dynamic quantization yields a 2.2× speedup (P50) and cuts RAM usage by 57.5%, top-line accuracy (92.16% → 91.03%) hides a severe drop in macro performance:

  • Macro F1 Collapse: Drops from 65.19% to 50.00%. Dynamic weight quantization compresses logit decision boundaries for rare token tags.
  • Punctuation Degradation: Punctuation F1 falls 10.66 percentage points (69.09% → 58.43%), causing increased missing or misclassified commas, colons, and sentence boundaries.
  • Casing Stability: Capitalization F1 remains mostly intact (91.71% → 89.38%).

Recommendation: Use ONNX FP32 for production pipelines where text formatting and punctuation precision are critical. Use ONNX INT8 in latency-critical environments where speed and memory constraints outweigh exact punctuation recovery.


📊 Benchmarks & Performance (Epoch 7 - Best Checkpoint)

Evaluated on a 95/5 train/holdout split across diverse Classical Latin literary and historical corpora.

Overall Summary Metrics

Metric Score
Overall Accuracy 92.16%
Macro F1 0.6523
Precision 62.28%
Recall 69.43%
Validation Loss 0.2592

🎯 Sub-Task Breakdown

Task Accuracy F1 Score
Casing Restoration 98.35% 0.9171
Punctuation Insertion 93.67% 0.6908

📈 Training Progression

The model was trained over 9 epochs fine-tuning latincy/latin-bert. Model weights from Epoch 7 were selected based on optimal overall F1.

Epoch Train Loss Val Loss Overall F1 Overall Acc Casing Acc Punct Acc
1 0.6284 0.2883 0.6097 91.05% 98.08% 92.79%
2 0.5617 0.2701 0.6275 91.60% 98.19% 93.26%
3 0.5150 0.2618 0.6388 91.91% 98.27% 93.50%
4 0.4855 0.2607 0.6455 92.04% 98.30% 93.60%
5 0.4652 0.2579 0.6440 92.12% 98.33% 93.65%
6 0.4502 0.2582 0.6462 92.15% 98.34% 93.66%
7 0.4351 0.2592 0.6523 92.16% 98.35% 93.67%
8 0.4262 0.2598 0.6495 92.24% 98.36% 93.75%
9 0.4172 0.2601 0.6507 92.21% 98.37% 93.71%

⚡ Hardware & Environmental Footprint

  • Hardware Infrastructure: NVIDIA L4 GPU via Modal
  • Training Time: 3.82 hours
  • Estimated Carbon Emissions: 0.1797 kg CO₂eq
Downloads last month
167
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for njand/latin-asr-postprocessor

Quantized
(1)
this model

Dataset used to train njand/latin-asr-postprocessor

Space using njand/latin-asr-postprocessor 1

Evaluation results

  • Macro F1 on Latin ASR Post-Processing Dataset
    self-reported
    0.652
  • Overall Accuracy on Latin ASR Post-Processing Dataset
    self-reported
    0.922
  • Precision on Latin ASR Post-Processing Dataset
    self-reported
    0.623
  • Recall on Latin ASR Post-Processing Dataset
    self-reported
    0.694