Spaces:
Sleeping
Sleeping
File size: 7,645 Bytes
8858e7b 51a27d1 8858e7b a90b0ef 8858e7b a90b0ef 8858e7b 51a27d1 d1aa2fe 51a27d1 d1aa2fe 51a27d1 7ec3e35 51a27d1 d1aa2fe 51a27d1 d1aa2fe 994eaff d1aa2fe 51a27d1 d1aa2fe 51a27d1 d1aa2fe 51a27d1 d1aa2fe 51a27d1 d1aa2fe 51a27d1 d1aa2fe 7ec3e35 d1aa2fe 7ec3e35 d1aa2fe 7ec3e35 d1aa2fe 7ec3e35 d1aa2fe 7ec3e35 d1aa2fe 51a27d1 d1aa2fe 389894e d1aa2fe 389894e d1aa2fe 7330a0e d1aa2fe 51a27d1 d1aa2fe 51a27d1 d1aa2fe 51a27d1 d1aa2fe 51a27d1 d1aa2fe 44b3467 d1aa2fe 44b3467 d1aa2fe 44b3467 d1aa2fe 44b3467 d1aa2fe 44b3467 d1aa2fe 7330a0e d1aa2fe 7330a0e d1aa2fe 7330a0e d1aa2fe 7330a0e d1aa2fe 7330a0e d1aa2fe 7330a0e d1aa2fe 44b3467 d1aa2fe 51a27d1 d1aa2fe 3decd42 d1aa2fe 3decd42 d1aa2fe 9ce0bb8 d1aa2fe 3decd42 d1aa2fe 3decd42 d1aa2fe 3decd42 d1aa2fe 3decd42 6a17342 d1aa2fe 3decd42 51a27d1 d1aa2fe 51a27d1 d1aa2fe 3decd42 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 | ---
title: BridgeLink ASL
sdk: gradio
sdk_version: 6.13.0
app_file: app.py
pinned: false
license: mit
python_version: 3.11
---
# BridgeLink ASL
BridgeLink ASL is a computer vision project for American Sign Language video understanding. The final project compares three paths:
- a word-level landmark CNN for live webcam sign recognition
- a sentence-level 3D CNN trained on How2Sign RGB clips
- an imported Qwen2.5-VL workspace for VLM-based sentence translation experiments
## Important clarification
- The final full-sentence classifier in this repo uses **How2Sign**, not WLASL-100.
- The sentence classifier is a **3D CNN over RGB video frames**, not the older 1D landmark CNN.
- The landmark CNN remains in the repo only for the **word-level live webcam demo** and isolated-sign experiments.
- If you are reviewing the final sentence-classification work, the main model to look at is the **How2Sign 3D CNN**.
## What the app does
The Gradio app supports three modes:
- **Live Webcam** - real-time word-level ASL recognition from MediaPipe Holistic landmarks.
- **Upload / Record Clip** - upload a short sign clip and run the word-level recognition branch.
- **How2Sign Sentence CNN** - upload a short RGB sentence clip and classify it with the closed-vocabulary 3D CNN trained on repeated How2Sign sentences.
For the final class project, the sentence branch is the key contribution because it operates on **full ASL sentence clips**, not still images and not isolated single-sign labels.
## Final sentence-classification path
The sentence classifier works as follows:
```text
How2Sign RGB sentence clip
-> decode video
-> uniformly sample 16 frames
-> resize frames to 112 x 112
-> stack into a 3D video volume
-> Conv3D + BatchNorm + MaxPool blocks
-> GlobalAveragePooling3D
-> Dense softmax classifier
-> best sentence label from the repeated-sentence How2Sign subset
```
Key details:
- Dataset: **How2Sign**
- Input: **16 RGB frames per clip**
- Frame size: **112 x 112**
- Model family: **3D CNN**
- Task: **closed-vocabulary sentence classification**
The final best sentence model is the normalized top-25 How2Sign repeated-sentence model, which reduces duplicate punctuation and wording variants into 21 usable sentence classes.
## Dataset summary
### How2Sign sentence data
How2Sign is the dataset used for the sentence classifier. It provides RGB videos aligned to English sentence translations. In this project, the 3D CNN uses short frontal-view RGB sentence clips.
Important limitation:
- The full local How2Sign inventory contains about **31k clips**, but most sentence labels appear only once.
- Because a standard softmax classifier needs repeated examples per class, the 3D CNN was trained on a **repeated-sentence subset** rather than treating all 31k clips as separate classes.
- This means the final 3D CNN is a **sentence classifier over repeated sentence labels**, not an open-ended translator over every unique How2Sign sentence.
### WLASL word-level data
WLASL-based data are kept only for the **word-level webcam / isolated-sign branch**. They are **not** the dataset used for the final sentence classifier.
## VLM workspace
The teammate Qwen2.5-VL fine-tuning workspace is included directly in this repo under `vlm_hf_space/`.
An additional organized copy is also present under `VLM/ASL-Video-To-Sentence-Translation/` on the current main branch.
That folder preserves:
- QLoRA training scripts
- How2Sign and ASL Citizen data-prep scripts
- experiment configs
- archived baseline vs fine-tuned metrics
- sample prediction outputs from the Hugging Face Space repo
Original source:
`https://huggingface.co/spaces/ofraij123/ASL-Video-To-Sentence-Translation`
## Run locally
```bash
py -3.11 -m venv .venv
# Windows
.venv\Scripts\Activate.ps1
# macOS / Linux
source .venv/bin/activate
pip install -r requirements.txt
python -m pip install -e .
python app.py
```
Open `http://127.0.0.1:7860` in your browser.
## Hugging Face Space configuration
For the word-level live demo:
- `HF_MODEL_REPO` should point to a model repo containing `cnn_landmark_wlasl25_best.pt`
- optionally set `HF_MODEL_FILENAME` if you publish that model under a different name
For the sentence-level How2Sign demo:
- `HF_SENTENCE_MODEL_REPO` should point to a model repo containing `cnn-3d-sentence-top25-normalized.keras`
- optionally set `HF_SENTENCE_MODEL_FILENAME` if you publish the sentence model under a different name
If model artifacts already live under `models/` in the Space repo, the app can load them locally without additional Hugging Face Hub variables.
## Training notes
### Word-level landmark branch
The older webcam branch uses:
- MediaPipe Holistic landmark extraction
- a rolling 32-frame sequence
- a temporal 1D CNN over landmark vectors
This branch is useful for the live demo, but it is **not** the final sentence-classification pipeline.
### Sentence-level 3D CNN branch
The main sentence training entrypoint is:
```bash
python -m bridgelink_asl.cli.train_cnn_model ^
--clips data\processed\how2sign_sentences_top25_normalized.frames.jsonl ^
--output models\cnn-3d-sentence-top25-normalized.keras ^
--batch-size 2 ^
--frame-count 16 ^
--image-size 112
```
This branch trains directly on **How2Sign sentence videos** and is the model that should be cited when describing the repo's final sentence classifier.
## Repo layout
```text
app.py Gradio and Hugging Face Space entrypoint
requirements.txt Space / local dependencies
src/bridgelink_asl/
inference.py Word-level landmark inference runtime
sentence_inference.py How2Sign sentence 3D CNN inference runtime
wrapper.py CNN / VLM comparison wrapper utilities
models/
cnn_landmark_best.pt Word-level landmark CNN weights
cnn_landmark_wlasl25_best.pt Smaller word-level live-demo weights
cnn-3d-sentence-top25-normalized.keras
Final How2Sign sentence 3D CNN weights
cnn-3d-sentence-top25-normalized.labels.json
Sentence label map
sign_transformer_best.pt Optional word-level attention experiment
results/
*.json Metrics and dataset summaries
*.png Plots for the report and presentation
notebooks/
train_wlasl100_colab.ipynb Word-level landmark training notebook
report/
main.tex Final report source
BridgeLink_ASL_Final_Report.pdf
Final compiled report
vlm_hf_space/ Imported Qwen2.5-VL workspace from Hugging Face Space
VLM/ASL-Video-To-Sentence-Translation/
Organized VLM workspace copy added on GitHub main
docs/ Setup notes, architecture notes, and planning docs
tests/ Unit tests
```
## Final project focus
To avoid confusion:
- **Word-level recognition** in this repo is the landmark CNN branch.
- **Full-sentence classification** in this repo is the **How2Sign 3D CNN** branch.
- **Open-ended sentence translation experiments** in this repo are in the imported **VLM workspace**.
If you need the final sentence model for the class project discussion, use the How2Sign 3D CNN path and not the older WLASL landmark branch.
## Team
| Name | Role |
|---|---|
| Dalen Gordon | Undergraduate |
| Ervin Gordon III | Undergraduate |
| Frank Garcia | Graduate |
| Omar Fraij | Undergraduate |
ITCS 4152/5010 - Introduction to Computer Vision, Spring 2026
University of North Carolina at Charlotte
|