silentvsr-models / README.md
ebowwa's picture
Upload README.md with huggingface_hub
136c131 verified
|
Raw
History Blame Contribute Delete
9.81 kB
metadata
license: mit
tags:
  - visual-speech-recognition
  - lip-reading
  - coreml
  - auto-avsr
  - silentvsr
  - ctc
language: en
base_model:
  - AD1TEYA/lip-reading-model
  - aaahmet/silent-lip-reader-model
pipeline_tag: automatic-speech-recognition

SilentVSR Core ML models

Visual-speech (lip-reading) models for the SilentVSR macOS/iOS client. This is a public, gated release — the model card is visible to everyone, but weight downloads (base and personalization) require an approved access request.

  • base/ — the general (speaker-independent) Core ML models the app bundles.
  • personalization/ — an experimental, speaker-specific rank-8 LoRA personalization (the ANEtest test variant), included to demonstrate that on-device personalization works end-to-end. It is one individual's visual-speech data, which is why the repository is gated rather than open.

This is a derivative/mirror, not an independently trained model. See Provenance.


Provenance

This model is a clone/derivative of AD1TEYA/lip-reading-model, which is itself a re-host of the official Auto-AVSR visual-speech checkpoint. The weights are byte-for-byte identical to the upstream lineage (verified by SHA-256; see fingerprints below). Nothing here was trained from scratch by the maintainer of this repo.

Lineage:

Auto-AVSR  (Pingchuan Ma et al., github.com/mpc001/auto_avsr)
   │  vsr_trlrs2lrs3vox2avsp_base.pth   (3,291 h, ~20.3% LRS3 WER, ~250M params)
   ▼
AD1TEYA/lip-reading-model        (HF re-host; NO model card, NO license, NO eval)
   │  pytorch_model.pt
   ▼
aaahmet/silent-lip-reader-model  (HF mirror; MIT; "not trained by the re-host";
   │                              "mirrored from AD1TEYA/lip-reading-model")
   │  revision ff1e35515594eb4562c469b99d147cf6a3d2891f, pytorch_model.pt
   ▼
Open-Altergo pinned checkpoint   (byte-identical; fingerprinted 2026-07-28)
   │  SHA-256 fbf7cd70ff1c0e694b3030fb779dbb4570f04e4b841d62f9296c229e94878ddb
   ▼
SilentVSRCTC  (coremltools conversion → Core ML, cloud/modal/convert_coreml.py)
   │  base/SilentVSRCTC.mlpackage (FP32, flexible) +
   │  base/SilentVSRCTC_fixed105_fp16.mlpackage (FP16, fixed-105)
   ▼
personalization/  (speaker-specific rank-8 LoRA fine-tune on top of the base)

License: mit — following the documented aaahmet/silent-lip-reader-model card and the upstream Auto-AVSR code license. ⚠️ Caveat: AD1TEYA/lip-reading-model itself documents no license and no evaluation, and the exact upstream training split / checkpoint-averaging recipe is not fully established. Treat the MIT tagging as a best-effort inference from the documented mirror; confirm upstream terms before any public redistribution.

Snapshot note: The base weights in this repo are a snapshot of the mirror above, fingerprint-verified on 2026-07-28 against the official Google-Drive model-zoo file.


Repository contents

Path Size What it is
base/SilentVSRCTC.mlpackage ~388 MB Base model, FP32, flexible-length (8–250 frames). The working baseline.
base/SilentVSRCTC_fixed105_fp16.mlpackage ~362 MB Base model, FP16, fixed 105-frame window — a Neural-Engine placement candidate.
personalization/SilentVSRFrontend.mlmodelc ~354 MB Personalized visual frontend/encoder (compiled Core ML).
personalization/SilentVSRCTCProjection.mlmodelc ~8 MB Personalized encoder→decoder projection head.
personalization/SilentVSRCTCProjectionPersonalizedANEtest.mlmodelc ~8 MB Pre-merged static personalized projection (ANE-targeted test variant).
personalization/adapter-a-rank8.f32le 24 KB Rank-8 LoRA adapter A (raw float32 little-endian).
personalization/adapter-b-rank8-tiled.f32le 160 KB Rank-8 LoRA adapter B, tiled.

Base artifacts are .mlpackage (source); personalization artifacts are compiled .mlmodelc bundles. Integrity is content-pinned by SHA-256 in the app manifest, not by filename.


Architecture (Auto-AVSR, visual-only, ~250M params)

Stage Configuration
3D frontend Conv3D 1→64, kernel 5×7×7, stride 1×2×2
Spatial trunk ResNet-18-style [2,2,2,2], widths 64/128/256/512
Visual frontend output [B, T, 512] per frame
Encoder projection Linear 512→768
Encoder Conformer, 12 blocks, 12 heads, FFN 3072, conv kernel 31
CTC head 5,049 output classes
Decoder Transformer, 6 blocks, width 768, 12 heads, FFN 3072
Objective 0.1·CTC + 0.9·attention (label smoothing 0.1)
Modality Video-only (no audio path)

One encoder feature vector is produced per input frame (temporal stride 1 in the frontend). T is dynamic in the source checkpoint; the fixed 105-frame window is a deployment choice, not a model constraint.


Core ML I/O contract

Flexible FP32 Fixed-105 FP16
Input name input input
Input shape [-1, -1, 1, 88, 88] (flex. T) [1, 105, 1, 88, 88]
Input dtype float32 float16
Output name var_2595 var_2595
Output shape [-1, -1, -1] [-1, -1, -1]
Frames 8–250 105
Deployment target macOS 14+ / iOS 17+ iOS 18+ / macOS 15+ (NE candidate)
  • Frame rate: 25 FPS · window: 105 frames · stride: 90 frames
  • Decoding: CTC prefix-beam search, beam width 20, blank token ID 0
  • Vocabulary (vocab.json) SHA-256: 076b4f80e4ce1a14f18c94d6185f2688c0b1819b434301d5847b6fe345270bcd
  • Artifact SHA-256: flexible-FP32 8a52b695…727f5 · fixed-105-FP16 f142eadc…e7276

Preprocessing contract

  1. Decode video, normalize to exactly 25 FPS.
  2. Face/landmark detection (MediaPipe/RetinaFace), affine stabilization.
  3. Extract 96×96 RGB mouth-centered crop per frame (intermediate).
  4. Random/center crop 96→88, RGB→1 grayscale channel.
  5. Normalize: mean 0.421, std 0.165 (after /255).
  6. Network tensor: [B, T, 1, 88, 88] float32 (frontend transposes to [B, 1, T, 88, 88] internally).

Vocabulary

Immutable SentencePiece unigram-5000 tokenizer that must travel with every compatible model:

ID 0       = <blank>  (CTC)
IDs 1–5047 = SentencePiece units
ID 5048    = shared <sos>/<eos>
odim       = 5049
ignore ID  = -1

Transcripts are uppercased before tokenization. Do not swap the tokenizer, unit file, ID ordering, blank ID, or EOS ID — weights will still load but decode to garbage.


Personalization

personalization/ is a speaker-specific adaptation of the base model for a single speaker — it encodes that speaker's mouth geometry, articulation, and capture conditions. This upload is the experimental ANEtest test variant, published (gated) to demonstrate the on-device personalization path; it is not claimed to be a production-quality personal model.

  • Method: rank-8 LoRA-style fine-tune, default encoder-decoder strategy (last 2 of 12 Conformer blocks + encoder projection + Transformer decoder + CTC head trainable; visual frontend frozen).
  • Runtime paths: the app applies personalization either as a pre-merged static projection (SilentVSRCTCProjectionPersonalizedANEtest.mlmodelc, coremlStatic) or by merging the rank-8 LoRA adapters at inference via Accelerate (swiftAccelerate). Adapter ID 7875BF49-…-rank8-v1, rank 8.
  • Personalization artifacts retain their original internal Core ML metadata (generatedClassName, sqeakd.personalization=lora, sqeakd.personalization.rank=8); only the repository filenames were renamed to the SilentVSR* scheme.

Training/conversion references: training/ and pretraining/auto-avsr-checkpoints/ in the source project.


Performance

No new benchmark was produced for this snapshot. Reported numbers are upstream and informal:

  • Official Auto-AVSR vsr_trlrs2lrs3vox2avsp_base.pth: ~20.3% WER on LRS3 (clean benchmark, upstream evaluation pipeline).
  • aaahmet/silent-lip-reader-model card claims ~25–30% WER on clean, frontal, deliberately-articulated English — informal, no evaluation table.
  • These are not expected WERs for webcam video, silent mouthing, a single personalized speaker, mobile capture, streaming windows, or out-of-domain language.
  • No personalization WER has been established yet.

Intended use & limitations

  • Intended: on-device visual-speech transcription for a cooperative, front-facing speaker in the SilentVSR client; research/personal use.
  • Not intended: surveillance, covert capture, speaker identification across unknown subjects, or any deployment without consent.
  • Limitations: lip-reading is inherently ambiguous (many phonemes are visually identical); accuracy drops sharply with head angle, lighting, distance, and casual/ silent speech.

Citation


Disclaimer

This is a derivative re-host/conversion, not an independently trained model. The base weights are byte-identical to the Auto-AVSR / AD1TEYA / aaahmet lineage. The repository is gated — vet access requests carefully, since personalization/ is speaker-specific (one individual's visual-speech data). Confirm upstream licenses before any further redistribution.