Spaces:
Sleeping
A newer version of the Gradio SDK is available: 6.22.0
Final Project Workflow
This branch makes the final professor-facing comparison explicit:
trained landmark CNN vs zero-shot VLM reranking
The Transformer remains in the notebook as an optional attention/extra-credit experiment, not the main comparison.
1. Train From The Colab Notebook
Use notebooks/train_wlasl100_colab.ipynb.
The notebook expects:
/content/drive/MyDrive/BridgeLink-ASL/data/wlasl-processed.zip
Run the notebook top-to-bottom. If landmarks are already extracted, Step 7 is
cached and should report mostly Cached samples instead of redoing the slow
MediaPipe pass.
2. Primary CNN Outputs
The CNN section saves:
/content/drive/MyDrive/BridgeLink-ASL/models/cnn_landmark_best.pt
/content/drive/MyDrive/BridgeLink-ASL/results/cnn_metrics.json
/content/drive/MyDrive/BridgeLink-ASL/results/cnn_training_curves.png
/content/drive/MyDrive/BridgeLink-ASL/results/cnn_confusion_matrix.png
/content/drive/MyDrive/BridgeLink-ASL/results/cnn_classification_report.txt
These are the main model artifacts for the report.
3. Optional Transformer Outputs
The Transformer cells are kept after the CNN section. Run them only if time allows and report them as an attention-based extension.
/content/drive/MyDrive/BridgeLink-ASL/models/sign_transformer_best.pt
/content/drive/MyDrive/BridgeLink-ASL/results/metrics.json
4. Build The CNN/VLM Eval Set
4. Train The Smaller Live Demo CNN
The notebook also trains a smaller WLASL-25 checkpoint for the live demo:
/content/drive/MyDrive/BridgeLink-ASL/models/cnn_landmark_wlasl25_best.pt
/content/drive/MyDrive/BridgeLink-ASL/results/demo_wlasl25_metrics.json
/content/drive/MyDrive/BridgeLink-ASL/results/demo_wlasl25_training_curves.png
/content/drive/MyDrive/BridgeLink-ASL/results/demo_wlasl25_confusion_matrix.png
Use this checkpoint for Hugging Face Space because it is much more likely to behave well live than the full 100-class model.
5. Build The CNN/VLM Eval Set
The CNN/VLM notebook cell creates:
/content/drive/MyDrive/BridgeLink-ASL/vlm_eval_wlasl25_cnn/wlasl25_cnn_hybrid_eval.jsonl
/content/drive/MyDrive/BridgeLink-ASL/vlm_eval_wlasl25_cnn/clips/
Each row includes the true label, clip path, CNN top-1, CNN top-5 candidates, and a constrained prompt for the VLM.
6. Score The VLM Reranker
Copy the eval folder into:
data/vlm_eval_wlasl25_cnn/
Generate the review sheet and baseline metrics:
python scripts/evaluate_hybrid_vlm.py `
--manifest data/vlm_eval_wlasl25_cnn/wlasl25_cnn_hybrid_eval.jsonl `
--output-dir results/vlm_eval
Fill results/vlm_eval/vlm_review_template.csv with VLM choices, then rescore:
python scripts/evaluate_hybrid_vlm.py `
--manifest data/vlm_eval_wlasl25_cnn/wlasl25_cnn_hybrid_eval.jsonl `
--predictions results/vlm_eval/vlm_review_template.csv `
--output-dir results/vlm_eval
Report:
CNN top-1 accuracy
CNN top-5 coverage
VLM-reranked top-5 accuracy
7. Deploy The Space
Upload cnn_landmark_best.pt to a Hugging Face model repo and set Space
variables:
HF_MODEL_REPO=<username>/<model-repo>
HF_MODEL_FILENAME=cnn_landmark_wlasl25_best.pt
The Space can still load sign_transformer_best.pt if you set
HF_MODEL_FILENAME=sign_transformer_best.pt, but the final demo should use the
WLASL-25 CNN checkpoint unless the professor asks for the Transformer extension.