File size: 3,388 Bytes
a84cb5d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
---
language: en
license: cc-by-nc-4.0
pipeline_tag: automatic-speech-recognition
tags:
  - lip-reading
  - visual-speech-recognition
  - auto-avsr
  - lrs3
  - audio-visual
datasets:
  - Ainncy/LRS3
metrics:
  - wer
model-index:
  - name: LRS3 Visual-Only Lip Reader
    results: []
---

# LRS3 Visual-Only Lip Reading Model

**Visual-only English lip-reading model** based on the [Auto-AVSR](https://github.com/mpc001/auto_avsr) architecture
(commit `182b628`), fine-tuned on the Ainncy/LRS3 trainval split.

## Model Description

- **Architecture**: End-to-end (E2E) visual-only speech recognition, 250M parameters
- **Modality**: Video only (mouth ROI crops, no audio)
- **Frontend**: MediaPipe face detection + mouth ROI cropping at 16-second segments
- **Training**: 20 epochs on a single NVIDIA A100-SXM4-80GB, learning rate 0.001
- **Framework**: PyTorch 2.5.1 + PyTorch Lightning 2.5.0
- **Checkpoint**: `model_avg_10.pth` — average of the last 10 epoch checkpoints

## Dataset

Trained on the **supervised trainval split** from [Ainncy/LRS3](https://huggingface.co/datasets/Ainncy/LRS3):

- **Training samples**: ~98% of trainval (lrs3_train_fit.csv)
- **Validation samples**: ~2% of trainval (lrs3_train_val.csv)
- **Note**: The official LRS3 test set is **not** distributed by Ainncy/LRS3.
  Validation metrics are computed on a deterministic held-out subset of trainval
  and should **not** be compared to published LRS3 test-set results.

## Intended Use

This model is intended for **research purposes only** in visual speech recognition
(lip reading). It does **not** use audio input and is designed to explore the
limits of visual-only speech recognition.

## Limitations

- **Visual-only**: Performance is inherently lower than audio-visual or audio-only models.
- **English only**: Trained solely on English speech from LRS3 (TED/TEDx talks).
- **Controlled environment**: LRS3 consists of frontal-facing speakers with relatively
  stable lighting. Real-world performance will be significantly lower.
- **No speaker diarization**: Cannot distinguish between multiple speakers.
- **Vocabulary constrained**: Limited to vocabulary present in the training data.
- **Privacy & ethics**: Lip-reading technology raises privacy concerns. This model
  should not be used for surveillance or non-consensual speech recognition.

## Training Setup

Preprocessing was performed on Modal cloud GPUs using the pipeline in the
[accompanying repository](https://github.com/yoannjardin78/lrs3-lipreader):

1. Mouth ROI detection with MediaPipe (CPU, 8 parallel shards)
2. Video cropping to 16-second segments
3. Training on 1× A100-80GB for 20 epochs (~4h40)

The raw MP4 video files, crops, and transcriptions are **not** included in this
repository and must not be redistributed without explicit authorization from the
dataset copyright holders.

## Citation

```bibtex
@article{ma2023auto,
  title={Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels},
  author={Ma, Pingchuan and Haliassos, Alexandros and Fernandez-Lopez, Adriana
          and Chen, Honglie and Petridis, Stavros and Pantic, Maja},
  journal={arXiv preprint arXiv:2303.08807},
  year={2023}
}
```

## License

This model is released under the Creative Commons Attribution-NonCommercial 4.0
International License (CC BY-NC 4.0), consistent with the LRS3 dataset terms.
Commercial use is prohibited.