silentvsr-models / README.md
ebowwa's picture
Upload README.md with huggingface_hub
136c131 verified
|
Raw
History Blame Contribute Delete
9.81 kB
---
license: mit
tags:
- visual-speech-recognition
- lip-reading
- coreml
- auto-avsr
- silentvsr
- ctc
language: en
base_model:
- AD1TEYA/lip-reading-model
- aaahmet/silent-lip-reader-model
pipeline_tag: automatic-speech-recognition
---
# SilentVSR Core ML models
Visual-speech (lip-reading) models for the **SilentVSR** macOS/iOS client.
This is a **public, gated** release — the model card is visible to everyone, but
weight downloads (base and personalization) require an approved access request.
- **`base/`** — the general (speaker-independent) Core ML models the app bundles.
- **`personalization/`** — an **experimental, speaker-specific** rank-8 LoRA
personalization (the `ANEtest` test variant), included to demonstrate that on-device
personalization works end-to-end. It is one individual's visual-speech data, which is
why the repository is gated rather than open.
> This is a **derivative/mirror**, not an independently trained model. See
> [Provenance](#provenance).
---
## Provenance
This model is a **clone/derivative of `AD1TEYA/lip-reading-model`**, which is itself
a re-host of the official **Auto-AVSR** visual-speech checkpoint. The weights are
**byte-for-byte identical** to the upstream lineage (verified by SHA-256; see
fingerprints below). Nothing here was trained from scratch by the maintainer of this
repo.
**Lineage:**
```
Auto-AVSR (Pingchuan Ma et al., github.com/mpc001/auto_avsr)
│ vsr_trlrs2lrs3vox2avsp_base.pth (3,291 h, ~20.3% LRS3 WER, ~250M params)
AD1TEYA/lip-reading-model (HF re-host; NO model card, NO license, NO eval)
│ pytorch_model.pt
aaahmet/silent-lip-reader-model (HF mirror; MIT; "not trained by the re-host";
│ "mirrored from AD1TEYA/lip-reading-model")
│ revision ff1e35515594eb4562c469b99d147cf6a3d2891f, pytorch_model.pt
Open-Altergo pinned checkpoint (byte-identical; fingerprinted 2026-07-28)
│ SHA-256 fbf7cd70ff1c0e694b3030fb779dbb4570f04e4b841d62f9296c229e94878ddb
SilentVSRCTC (coremltools conversion → Core ML, cloud/modal/convert_coreml.py)
│ base/SilentVSRCTC.mlpackage (FP32, flexible) +
│ base/SilentVSRCTC_fixed105_fp16.mlpackage (FP16, fixed-105)
personalization/ (speaker-specific rank-8 LoRA fine-tune on top of the base)
```
**License:** `mit` — following the documented `aaahmet/silent-lip-reader-model` card
and the upstream Auto-AVSR code license. ⚠️ Caveat: `AD1TEYA/lip-reading-model`
itself documents **no license and no evaluation**, and the exact upstream training
split / checkpoint-averaging recipe is not fully established. Treat the MIT tagging
as a best-effort inference from the documented mirror; confirm upstream terms before
any **public** redistribution.
**Snapshot note:** The base weights in this repo are a snapshot of the mirror above,
fingerprint-verified on 2026-07-28 against the official Google-Drive model-zoo file.
---
## Repository contents
| Path | Size | What it is |
|---|---:|---|
| `base/SilentVSRCTC.mlpackage` | ~388 MB | Base model, **FP32**, flexible-length (8–250 frames). The working baseline. |
| `base/SilentVSRCTC_fixed105_fp16.mlpackage` | ~362 MB | Base model, **FP16**, fixed 105-frame window — a Neural-Engine placement candidate. |
| `personalization/SilentVSRFrontend.mlmodelc` | ~354 MB | Personalized **visual frontend/encoder** (compiled Core ML). |
| `personalization/SilentVSRCTCProjection.mlmodelc` | ~8 MB | Personalized **encoder→decoder projection** head. |
| `personalization/SilentVSRCTCProjectionPersonalizedANEtest.mlmodelc` | ~8 MB | Pre-merged **static** personalized projection (ANE-targeted test variant). |
| `personalization/adapter-a-rank8.f32le` | 24 KB | Rank-8 LoRA adapter A (raw float32 little-endian). |
| `personalization/adapter-b-rank8-tiled.f32le` | 160 KB | Rank-8 LoRA adapter B, tiled. |
Base artifacts are `.mlpackage` (source); personalization artifacts are compiled
`.mlmodelc` bundles. Integrity is content-pinned by SHA-256 in the app manifest, not
by filename.
---
## Architecture (Auto-AVSR, visual-only, ~250M params)
| Stage | Configuration |
|---|---|
| 3D frontend | Conv3D `1→64`, kernel `5×7×7`, stride `1×2×2` |
| Spatial trunk | ResNet-18-style `[2,2,2,2]`, widths 64/128/256/512 |
| Visual frontend output | `[B, T, 512]` per frame |
| Encoder projection | Linear `512→768` |
| Encoder | **Conformer**, 12 blocks, 12 heads, FFN 3072, conv kernel 31 |
| CTC head | 5,049 output classes |
| Decoder | **Transformer**, 6 blocks, width 768, 12 heads, FFN 3072 |
| Objective | `0.1·CTC + 0.9·attention` (label smoothing 0.1) |
| Modality | **Video-only** (no audio path) |
One encoder feature vector is produced per input frame (temporal stride 1 in the
frontend). `T` is dynamic in the source checkpoint; the fixed 105-frame window is a
**deployment** choice, not a model constraint.
---
## Core ML I/O contract
| | Flexible FP32 | Fixed-105 FP16 |
|---|---|---|
| Input name | `input` | `input` |
| Input shape | `[-1, -1, 1, 88, 88]` (flex. T) | `[1, 105, 1, 88, 88]` |
| Input dtype | float32 | float16 |
| Output name | `var_2595` | `var_2595` |
| Output shape | `[-1, -1, -1]` | `[-1, -1, -1]` |
| Frames | 8–250 | 105 |
| Deployment target | macOS 14+ / iOS 17+ | iOS 18+ / macOS 15+ (NE candidate) |
- **Frame rate:** 25 FPS · **window:** 105 frames · **stride:** 90 frames
- **Decoding:** CTC prefix-beam search, beam width 20, blank token ID 0
- **Vocabulary (`vocab.json`)** SHA-256: `076b4f80e4ce1a14f18c94d6185f2688c0b1819b434301d5847b6fe345270bcd`
- **Artifact SHA-256:** flexible-FP32 `8a52b695…727f5` · fixed-105-FP16 `f142eadc…e7276`
---
## Preprocessing contract
1. Decode video, normalize to **exactly 25 FPS**.
2. Face/landmark detection (MediaPipe/RetinaFace), affine stabilization.
3. Extract **96×96** RGB mouth-centered crop per frame (intermediate).
4. Random/center crop **96→88**, RGB→**1 grayscale channel**.
5. Normalize: **mean `0.421`, std `0.165`** (after `/255`).
6. Network tensor: **`[B, T, 1, 88, 88]` float32** (frontend transposes to
`[B, 1, T, 88, 88]` internally).
---
## Vocabulary
Immutable SentencePiece **unigram-5000** tokenizer that must travel with every
compatible model:
```
ID 0 = <blank> (CTC)
IDs 1–5047 = SentencePiece units
ID 5048 = shared <sos>/<eos>
odim = 5049
ignore ID = -1
```
Transcripts are **uppercased** before tokenization. Do not swap the tokenizer, unit
file, ID ordering, blank ID, or EOS ID — weights will still load but decode to garbage.
---
## Personalization
`personalization/` is a **speaker-specific** adaptation of the base model for a single
speaker — it encodes that speaker's mouth geometry, articulation, and capture
conditions. This upload is the **experimental `ANEtest` test variant**, published
(gated) to demonstrate the on-device personalization path; it is not claimed to be a
production-quality personal model.
- **Method:** rank-8 LoRA-style fine-tune, default `encoder-decoder` strategy
(last 2 of 12 Conformer blocks + encoder projection + Transformer decoder + CTC head
trainable; visual frontend frozen).
- **Runtime paths:** the app applies personalization either as a pre-merged **static
projection** (`SilentVSRCTCProjectionPersonalizedANEtest.mlmodelc`, `coremlStatic`)
or by merging the rank-8 **LoRA adapters** at inference via Accelerate
(`swiftAccelerate`). Adapter ID `7875BF49-…-rank8-v1`, rank 8.
- Personalization artifacts retain their original internal Core ML metadata
(`generatedClassName`, `sqeakd.personalization=lora`, `sqeakd.personalization.rank=8`);
only the repository filenames were renamed to the `SilentVSR*` scheme.
Training/conversion references: `training/` and `pretraining/auto-avsr-checkpoints/`
in the source project.
---
## Performance
**No new benchmark was produced for this snapshot.** Reported numbers are upstream and
informal:
- Official **Auto-AVSR `vsr_trlrs2lrs3vox2avsp_base.pth`: ~20.3% WER on LRS3** (clean
benchmark, upstream evaluation pipeline).
- `aaahmet/silent-lip-reader-model` card claims **~25–30% WER on clean, frontal,
deliberately-articulated English** — informal, no evaluation table.
- These are **not** expected WERs for webcam video, silent mouthing, a single
personalized speaker, mobile capture, streaming windows, or out-of-domain language.
- No personalization WER has been established yet.
---
## Intended use & limitations
- **Intended:** on-device visual-speech transcription for a cooperative, front-facing
speaker in the SilentVSR client; research/personal use.
- **Not intended:** surveillance, covert capture, speaker identification across
unknown subjects, or any deployment without consent.
- **Limitations:** lip-reading is inherently ambiguous (many phonemes are visually
identical); accuracy drops sharply with head angle, lighting, distance, and casual/
silent speech.
---
## Citation
- **Auto-AVSR** — Pingchuan Ma et al., *"Auto-AVSR: Audio-Visual Speech Recognition
with Automatic Labels."* Code: https://github.com/mpc001/auto_avsr
- **`AD1TEYA/lip-reading-model`** — https://huggingface.co/AD1TEYA/lip-reading-model
(re-host; undocumented)
- **`aaahmet/silent-lip-reader-model`** — Ahmet Dedeler,
https://huggingface.co/aaahmet/silent-lip-reader-model (MIT mirror; documented)
---
## Disclaimer
This is a **derivative re-host/conversion**, not an independently trained model. The
base weights are byte-identical to the Auto-AVSR / AD1TEYA / aaahmet lineage. The
repository is gated — vet access requests carefully, since `personalization/` is
speaker-specific (one individual's visual-speech data). Confirm upstream licenses
before any further redistribution.