--- language: en license: cc-by-nc-4.0 pipeline_tag: automatic-speech-recognition tags: - lip-reading - visual-speech-recognition - auto-avsr - lrs3 - audio-visual datasets: - Ainncy/LRS3 metrics: - wer model-index: - name: LRS3 Visual-Only Lip Reader results: [] --- # LRS3 Visual-Only Lip Reading Model **Visual-only English lip-reading model** based on the [Auto-AVSR](https://github.com/mpc001/auto_avsr) architecture (commit `182b628`), fine-tuned on the Ainncy/LRS3 trainval split. ## Model Description - **Architecture**: End-to-end (E2E) visual-only speech recognition, 250M parameters - **Modality**: Video only (mouth ROI crops, no audio) - **Frontend**: MediaPipe face detection + mouth ROI cropping at 16-second segments - **Training**: 20 epochs on a single NVIDIA A100-SXM4-80GB, learning rate 0.001 - **Framework**: PyTorch 2.5.1 + PyTorch Lightning 2.5.0 - **Checkpoint**: `model_avg_10.pth` — average of the last 10 epoch checkpoints ## Dataset Trained on the **supervised trainval split** from [Ainncy/LRS3](https://huggingface.co/datasets/Ainncy/LRS3): - **Training samples**: ~98% of trainval (lrs3_train_fit.csv) - **Validation samples**: ~2% of trainval (lrs3_train_val.csv) - **Note**: The official LRS3 test set is **not** distributed by Ainncy/LRS3. Validation metrics are computed on a deterministic held-out subset of trainval and should **not** be compared to published LRS3 test-set results. ## Intended Use This model is intended for **research purposes only** in visual speech recognition (lip reading). It does **not** use audio input and is designed to explore the limits of visual-only speech recognition. ## Limitations - **Visual-only**: Performance is inherently lower than audio-visual or audio-only models. - **English only**: Trained solely on English speech from LRS3 (TED/TEDx talks). - **Controlled environment**: LRS3 consists of frontal-facing speakers with relatively stable lighting. Real-world performance will be significantly lower. - **No speaker diarization**: Cannot distinguish between multiple speakers. - **Vocabulary constrained**: Limited to vocabulary present in the training data. - **Privacy & ethics**: Lip-reading technology raises privacy concerns. This model should not be used for surveillance or non-consensual speech recognition. ## Training Setup Preprocessing was performed on Modal cloud GPUs using the pipeline in the [accompanying repository](https://github.com/yoannjardin78/lrs3-lipreader): 1. Mouth ROI detection with MediaPipe (CPU, 8 parallel shards) 2. Video cropping to 16-second segments 3. Training on 1× A100-80GB for 20 epochs (~4h40) The raw MP4 video files, crops, and transcriptions are **not** included in this repository and must not be redistributed without explicit authorization from the dataset copyright holders. ## Citation ```bibtex @article{ma2023auto, title={Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels}, author={Ma, Pingchuan and Haliassos, Alexandros and Fernandez-Lopez, Adriana and Chen, Honglie and Petridis, Stavros and Pantic, Maja}, journal={arXiv preprint arXiv:2303.08807}, year={2023} } ``` ## License This model is released under the Creative Commons Attribution-NonCommercial 4.0 International License (CC BY-NC 4.0), consistent with the LRS3 dataset terms. Commercial use is prohibited.