BridgeLinkASL / README.md
ofraij123's picture
Sync from GitHub via hub-sync
d1aa2fe verified
|
Raw
History Blame Contribute Delete
7.65 kB
---
title: BridgeLink ASL
sdk: gradio
sdk_version: 6.13.0
app_file: app.py
pinned: false
license: mit
python_version: 3.11
---
# BridgeLink ASL
BridgeLink ASL is a computer vision project for American Sign Language video understanding. The final project compares three paths:
- a word-level landmark CNN for live webcam sign recognition
- a sentence-level 3D CNN trained on How2Sign RGB clips
- an imported Qwen2.5-VL workspace for VLM-based sentence translation experiments
## Important clarification
- The final full-sentence classifier in this repo uses **How2Sign**, not WLASL-100.
- The sentence classifier is a **3D CNN over RGB video frames**, not the older 1D landmark CNN.
- The landmark CNN remains in the repo only for the **word-level live webcam demo** and isolated-sign experiments.
- If you are reviewing the final sentence-classification work, the main model to look at is the **How2Sign 3D CNN**.
## What the app does
The Gradio app supports three modes:
- **Live Webcam** - real-time word-level ASL recognition from MediaPipe Holistic landmarks.
- **Upload / Record Clip** - upload a short sign clip and run the word-level recognition branch.
- **How2Sign Sentence CNN** - upload a short RGB sentence clip and classify it with the closed-vocabulary 3D CNN trained on repeated How2Sign sentences.
For the final class project, the sentence branch is the key contribution because it operates on **full ASL sentence clips**, not still images and not isolated single-sign labels.
## Final sentence-classification path
The sentence classifier works as follows:
```text
How2Sign RGB sentence clip
-> decode video
-> uniformly sample 16 frames
-> resize frames to 112 x 112
-> stack into a 3D video volume
-> Conv3D + BatchNorm + MaxPool blocks
-> GlobalAveragePooling3D
-> Dense softmax classifier
-> best sentence label from the repeated-sentence How2Sign subset
```
Key details:
- Dataset: **How2Sign**
- Input: **16 RGB frames per clip**
- Frame size: **112 x 112**
- Model family: **3D CNN**
- Task: **closed-vocabulary sentence classification**
The final best sentence model is the normalized top-25 How2Sign repeated-sentence model, which reduces duplicate punctuation and wording variants into 21 usable sentence classes.
## Dataset summary
### How2Sign sentence data
How2Sign is the dataset used for the sentence classifier. It provides RGB videos aligned to English sentence translations. In this project, the 3D CNN uses short frontal-view RGB sentence clips.
Important limitation:
- The full local How2Sign inventory contains about **31k clips**, but most sentence labels appear only once.
- Because a standard softmax classifier needs repeated examples per class, the 3D CNN was trained on a **repeated-sentence subset** rather than treating all 31k clips as separate classes.
- This means the final 3D CNN is a **sentence classifier over repeated sentence labels**, not an open-ended translator over every unique How2Sign sentence.
### WLASL word-level data
WLASL-based data are kept only for the **word-level webcam / isolated-sign branch**. They are **not** the dataset used for the final sentence classifier.
## VLM workspace
The teammate Qwen2.5-VL fine-tuning workspace is included directly in this repo under `vlm_hf_space/`.
An additional organized copy is also present under `VLM/ASL-Video-To-Sentence-Translation/` on the current main branch.
That folder preserves:
- QLoRA training scripts
- How2Sign and ASL Citizen data-prep scripts
- experiment configs
- archived baseline vs fine-tuned metrics
- sample prediction outputs from the Hugging Face Space repo
Original source:
`https://huggingface.co/spaces/ofraij123/ASL-Video-To-Sentence-Translation`
## Run locally
```bash
py -3.11 -m venv .venv
# Windows
.venv\Scripts\Activate.ps1
# macOS / Linux
source .venv/bin/activate
pip install -r requirements.txt
python -m pip install -e .
python app.py
```
Open `http://127.0.0.1:7860` in your browser.
## Hugging Face Space configuration
For the word-level live demo:
- `HF_MODEL_REPO` should point to a model repo containing `cnn_landmark_wlasl25_best.pt`
- optionally set `HF_MODEL_FILENAME` if you publish that model under a different name
For the sentence-level How2Sign demo:
- `HF_SENTENCE_MODEL_REPO` should point to a model repo containing `cnn-3d-sentence-top25-normalized.keras`
- optionally set `HF_SENTENCE_MODEL_FILENAME` if you publish the sentence model under a different name
If model artifacts already live under `models/` in the Space repo, the app can load them locally without additional Hugging Face Hub variables.
## Training notes
### Word-level landmark branch
The older webcam branch uses:
- MediaPipe Holistic landmark extraction
- a rolling 32-frame sequence
- a temporal 1D CNN over landmark vectors
This branch is useful for the live demo, but it is **not** the final sentence-classification pipeline.
### Sentence-level 3D CNN branch
The main sentence training entrypoint is:
```bash
python -m bridgelink_asl.cli.train_cnn_model ^
--clips data\processed\how2sign_sentences_top25_normalized.frames.jsonl ^
--output models\cnn-3d-sentence-top25-normalized.keras ^
--batch-size 2 ^
--frame-count 16 ^
--image-size 112
```
This branch trains directly on **How2Sign sentence videos** and is the model that should be cited when describing the repo's final sentence classifier.
## Repo layout
```text
app.py Gradio and Hugging Face Space entrypoint
requirements.txt Space / local dependencies
src/bridgelink_asl/
inference.py Word-level landmark inference runtime
sentence_inference.py How2Sign sentence 3D CNN inference runtime
wrapper.py CNN / VLM comparison wrapper utilities
models/
cnn_landmark_best.pt Word-level landmark CNN weights
cnn_landmark_wlasl25_best.pt Smaller word-level live-demo weights
cnn-3d-sentence-top25-normalized.keras
Final How2Sign sentence 3D CNN weights
cnn-3d-sentence-top25-normalized.labels.json
Sentence label map
sign_transformer_best.pt Optional word-level attention experiment
results/
*.json Metrics and dataset summaries
*.png Plots for the report and presentation
notebooks/
train_wlasl100_colab.ipynb Word-level landmark training notebook
report/
main.tex Final report source
BridgeLink_ASL_Final_Report.pdf
Final compiled report
vlm_hf_space/ Imported Qwen2.5-VL workspace from Hugging Face Space
VLM/ASL-Video-To-Sentence-Translation/
Organized VLM workspace copy added on GitHub main
docs/ Setup notes, architecture notes, and planning docs
tests/ Unit tests
```
## Final project focus
To avoid confusion:
- **Word-level recognition** in this repo is the landmark CNN branch.
- **Full-sentence classification** in this repo is the **How2Sign 3D CNN** branch.
- **Open-ended sentence translation experiments** in this repo are in the imported **VLM workspace**.
If you need the final sentence model for the class project discussion, use the How2Sign 3D CNN path and not the older WLASL landmark branch.
## Team
| Name | Role |
|---|---|
| Dalen Gordon | Undergraduate |
| Ervin Gordon III | Undergraduate |
| Frank Garcia | Graduate |
| Omar Fraij | Undergraduate |
ITCS 4152/5010 - Introduction to Computer Vision, Spring 2026
University of North Carolina at Charlotte