LSF_transcription / README.md
SamDNX's picture
Add 2nd Elix video per word (143 samples) + one-command run.sh
3b55737
|
Raw
History Blame Contribute Delete
10.2 kB
metadata
title: LSF Interpreter
emoji: 🤟
colorFrom: green
colorTo: blue
sdk: docker
app_port: 7860
pinned: false

DISCLAIMER: Used claude for django, UI, Gunicorn and some comments (and maybe git)

LSF_transcripter — French Sign Language interpreter

Real-time Langue des Signes Française (LSF) interpreter. It reads the signer from a camera, tracks the body parts that carry meaning in sign language — fingers / hand shape, eyes & gaze, mouth, chest and arms — draws the skeleton over the live feed, and transcribes recognised signs underneath.

A Django web UI shows the live camera feed with the skeleton at the top and the running transcription at the bottom (HTML / CSS / JS in separate files).

Architecture: the browser captures the camera and posts frames; the server runs MediaPipe + the recognizer and returns landmarks + glosses that the browser draws. The same app therefore runs locally and on Hugging Face (where the server has no webcam), and dual-pixel detection reads the real client device. See Deploy.


How it works

camera frame
   │
   ├─▶ MediaPipe Holistic ──▶ pose (33) · face mesh (468) · 2×hands (21)
   │         │                     fingers ▸ eyes ▸ mouth ▸ chest ▸ arms
   │         ▼
   │   depth refinement  ◀── dual-pixel / depth stream  (or MiDaS, or none)
   │         ▼
   │   normalised body-relative features  (features.py)
   │         ▼
   │   sign recogniser  ──▶  DTW templates · trained model · heuristics
   │         ▼
   └─▶ skeleton overlay  +  transcribed gloss
Body part the brief asked for Where it comes from
Fingers / hand shape left_hand / right_hand (21 pts each)
Eyes / gaze face-mesh eye + iris landmarks
Mouth (mouthing, non-manuals) face-mesh lip landmarks
Chest pose shoulders + hips
Arms pose shoulders / elbows / wrists
Hand hand wrist + pose wrist

Dual-pixel sensor detection (e.g. Pixel devices)

"Dual-pixel" depth is produced by the camera stack on certain devices (e.g. Pixel phones); it lives on the client device, not the server. Detection therefore runs in the browser (app.js, detectDualPixel): it probes the MediaStreamTrack capabilities (getCapabilities, getSupportedConstraints), enumerates devices for a depth camera, and matches the camera label (pixel, depth, tof, truedepth). The result is POSTed to /depth; the status pill then shows dual_pixel vs mediapipe, with the reason in its tooltip.

Honest limit: browsers don't generally expose the dual-pixel depth map itself, so when detected we flag the source and keep MediaPipe's relative z. If a platform does hand you a depth map, feed it to DepthEstimator.depth_for_frame(rgb, external_depth=...) and every landmark's z is resampled from it. A server-side MiDaS monocular fallback is also available (enable_monocular_depth, needs torch).

Recognition — be realistic

There is no off-the-shelf production LSF translation model. This project gives you a working, extensible pipeline rather than a magic black box:

  1. Vocabulary catalogue. It ships with ~80 common LSF words (signs/lexicon.json) grouped by theme — greetings, family, verbs, questions, emotions, time, food, places, colours. These are the targets shown in the web UI's vocabulary browser; a word becomes recognisable once you train it (step 1). You can add your own words from the UI too.
  2. Learn-by-example (default). Click Entraîner next to a word, perform the sign, click Sauvegarder. Samples are stored as normalised landmark sequences and matched with Dynamic Time Warping. Runs on a laptop CPU, no training needed, good for a small/medium vocabulary.
  3. Trained model (optional). Drop a Keras models/model.h5 (+ models/model.labels.json) and it's used automatically. Train your own sequence model on a corpus for real coverage.
  4. Heuristics. A couple of geometry-only signs so the demo isn't silent before you've recorded anything.

Setup

./setup.sh                       # creates .venv and installs deps
source .venv/bin/activate

Requires Python 3.9–3.11 (MediaPipe constraint). On macOS, grant camera access to your terminal in System Settings ▸ Privacy & Security ▸ Camera.

Run

Quick start — one command (creates the venv on first run):

./run.sh            # → http://127.0.0.1:8000   (ships pre-trained, 72 LSF signs)

Web UI (Django) — everything is in the browser, no terminal needed:

python web/manage.py runserver
# open http://127.0.0.1:8000  (allow the camera when prompted)

The browser captures the camera; getUserMedia needs a secure context, which localhost satisfies (on a LAN IP it won't — use localhost or HTTPS).

  • Top-left: live feed with the skeleton (drawn in the browser).
  • Bottom-left: transcription.
  • Right: vocabulary browser — search/filter the ~80-word catalogue, see which signs are trained, Entraîner a word (record → Sauvegarder), delete trained samples, Ajouter a custom word, or click a word for its tip + Elix sign video.
  • Status pills show the depth source (incl. dual_pixel when detected), FPS, and how many signs are trained.

Standalone window (local OpenCV, uses the server's own webcam):

python main.py            # keys: r=record, s=save, c=clear, q=quit

Deploy — Docker / Hugging Face

Docker (works locally too — the camera is the browser's, no device passthrough):

docker build -t lsf .
docker run -p 7860:7860 lsf
# open http://localhost:7860

Hugging Face Spaces: push the repo to a Docker Space (the front-matter at the top of this README sets sdk: docker and app_port: 7860). HF builds the Dockerfile and serves it over HTTPS, so the browser camera and dual-pixel detection work. Tuning is via env vars: LSF_MODEL_COMPLEXITY (0–2), LSF_REFINE_FACE (0/1), LSF_DEBUG, LSF_SECRET_KEY.

Note: a hosted Space has an ephemeral, often read-only filesystem — trained templates/added words persist in memory for the session but may not survive a restart. For durable training, run locally or import clips (below) into a mounted volume.

Pre-training words from video clips (batch)

Rather than recording each sign live, you can batch-import reference clips so the words are recognised immediately. Name clips by their gloss:

clips/BONJOUR.mp4            # one sample for BONJOUR
clips/AU REVOIR/01.mp4       # several samples for AU REVOIR
clips/AU REVOIR/02.mp4
python tools/import_videos.py clips/
# or a single file:
python tools/import_videos.py merci.mov --gloss MERCI

This runs the clips through the same pipeline as the live recorder and writes templates into signs/lsf_signs.json. Record 2–3 clips per sign (slightly different speed/framing) for better accuracy. Note: there is no public drop-in pre-trained LSF model — reference clips (filmed by you or any LSF video you're allowed to use) are how the vocabulary gets "pre-trained".

Auto-train from the Elix dictionary. tools/elix_train.py fetches each catalogue word's reference sign video from dico.elix-lsf.fr and trains a template from it — 72 of the 81 words ship pre-trained this way:

python tools/elix_train.py            # all words   (or --words BONJOUR MERCI …)

It rate-limits and identifies itself; robots.txt permits crawling. The Elix videos are © Signes de sens — this derives abstract landmark templates for personal/on-device use and does not redistribute the videos. The 9 unmatched words are multi-word phrases (AU REVOIR, ÇA VA, …) — record those by hand.


Layout

main.py                     standalone CLI runner (OpenCV window, local webcam)
tools/import_videos.py      batch clips → templates ("pre-train" words)
Dockerfile / .dockerignore  container build (HF Spaces / local)
requirements.txt            deps (mediapipe, opencv, django, gunicorn, whitenoise)
setup.sh                    venv bootstrap
signs/lexicon.json          ~80-word LSF vocabulary catalogue (themes + tips)
signs/lsf_signs.json        learned-sign templates (shared by CLI + web)
models/                     optional model.h5 + model.labels.json
lsf/
  landmarks.py              MediaPipe Holistic wrapper + compact points payload
  skeleton.py               server-side skeleton draw (CLI only)
  depth.py                  depth refinement + MiDaS fallback
  features.py               landmarks ▸ normalised body-relative feature vector
  recognizer.py             motion segmentation + DTW / model / heuristics
  lexicon.py                loads/extends the vocabulary catalogue (+ Elix links)
  camera.py                 InterpreterPipeline + local-webcam singleton (CLI)
  session.py                FrameSession: browser-pushed frames → landmarks+glosses
web/
  manage.py
  web/                      Django project (settings, urls, wsgi, asgi)
  interpreter/              Django app
    views.py                page · /process · /depth · vocab + training APIs
    urls.py
    templates/interpreter/index.html
    static/interpreter/style.css
    static/interpreter/app.js   getUserMedia · dual-pixel · /process · skeleton draw

Limitations & next steps

  • Out of the box it recognises only what you record (plus 2 rough heuristics). For broad LSF coverage, collect a dataset and train the model hook.
  • LSF grammar (spatial reference, classifiers, facial grammar) is not parsed — output is a stream of glosses, not fluent French. A gloss→French language step would sit after the recogniser.
  • Dual-pixel depth needs hardware that exposes it (see above).