Spaces:
Sleeping
A newer version of the Gradio SDK is available: 6.22.0
BridgeLink ASL Updated Roadmap
Current State
The repo currently has a working word-level architecture:
- webcam or synthetic frame source
- frame-to-landmark boundary
- baseline classifier that emits one sign label at a time
- smoothing to avoid repeated noisy labels
- text translation and speech fallback
- data collection support for isolated sign landmark samples
This is a good baseline, but it does not yet satisfy the professor-requested model comparison. The next version needs two comparable paths over the same sentence clips: a CNN baseline and a VLM interpretation path.
The LSTM/landmark path is now a possible stretch or support tool, not the core comparison. For the graded comparison, the two headline models are:
- CNN baseline: sampled video frames from a clip -> sentence/gloss class
- VLM path: sampled frames plus optional token trace -> natural English sentence
Local VLM Choice
The planned 30B-class VLM is Qwen/Qwen2.5-VL-32B-Instruct-AWQ.
Reasons:
- it is a practical 32B AWQ quantized Qwen2.5-VL target
- it supports image and video-style inputs
- it has explicit long-video and temporal event-understanding support
- it can produce structured outputs, which we need for parseable sentence events
- it can accept sampled frames plus a structured token trace from our current classifier
This is still a high-hardware model. If the team machine or Hugging Face Space cannot run the 32B AWQ model smoothly, use Qwen/Qwen2.5-VL-7B-Instruct as the emergency lower-hardware fallback. If stronger hardware is available, the original Qwen/Qwen2.5-VL-72B-Instruct target can remain a stretch comparison. The project must also keep a mock VLM provider for tests and offline demos.
We are not planning to depend on a cloud model for the final demo. The local model should be downloaded ahead of time, and the demo should still have a no-network fallback.
Updated Four Phases
- Stabilize the current repo and word-level baseline.
- Build sentence-window data and the CNN baseline.
- Add the VLM sentence interpreter and comparison wrapper.
- Harden the final demo and evaluation story.
The detailed phase docs live in docs/phases.
Target Architecture
camera frames
-> sentence window builder
-> sampled keyframes
-> CNN baseline OR VLM sentence interpreter
-> comparable model result
-> transcript and speech output
The CNN and VLM must read the same clip manifest so accuracy, latency, and failure cases can be compared fairly. The existing word classifier can still provide a token trace as grounding for the VLM, but it is not the main model being compared.
Demo Target
The final demo should support two flows:
wordmode: current stable baseline that emits recognized signscnnmode: short clip/window input classified by the sampled-frame CNNvlmmode: short clip/window input interpreted by Qwen2.5-VLcomparemode: run CNN and VLM on the same clip manifest and report accuracy/latency/failure modeswebcam clipmode: record a short live clip in the Gradio Space, then run the same compare flow
The backup demo should use pre-recorded sentence clips if live signing or API access fails.
Hosted Proof-Of-Concept Status
The Hugging Face Space branch now supports a complete hosted PoC flow plus dataset/experiment/report dashboard tabs. It does not require a trained CNN artifact or real Qwen runtime to boot. Instead, it uses video features and grounded fallback outputs while preserving the same result shape that the real models will use.
This means the project is possible as a Space-hosted final product if the final scope is a proof of concept. A fully accurate production version depends on dataset collection, CNN training, and GPU-backed VLM integration.