luciayen's picture
Upload best AfriSign Encoder checkpoints and documentation
3d02762 verified
|
Raw
History Blame Contribute Delete
2.84 kB
---
library_name: pytorch
license: other
tags:
- sign-language-recognition
- african-sign-language
- multimodal
- pose-estimation
- rgb-video
- metric-learning
- ksl
- casl
- nsl
datasets:
- mwakalucky/casl-w60
- awsaf49/kslc-dataset
metrics:
- accuracy
- f1
---
# AfriSign Encoder Best Models
This repository contains the best focused AfriSign Encoder checkpoints for the
KSL + CASL + NSL contribution. The models are trained for African sign-language
recognition using shared multilingual representations across pose landmarks,
RGB video frames, and static sign images.
## What Is Included
| Folder | Model | Checkpoint | Result file |
|---|---|---|---|
| `models/e9_5_research_v2_pose` | AfriSign Encoder E9.5 Research-v2 Pose-only | `pytorch_model.pt` | `results.json` |
| `models/e9_6_research_v2_rgb_image` | AfriSign Encoder E9.6 Research-v2 RGB/image-only | `pytorch_model.pt` | `results.json` |
| `models/e9_7_research_v2_multimodal` | AfriSign Encoder E9.7 Research-v2 RGB+pose/image | `pytorch_model.pt` | `results.json` |
## Main Results
| Experiment | Model | Modality | Test macro-F1 | Note |
|---|---|---:|---:|---|
| E9.1 | naive pooled pose-only | pose | 32.00% | baseline |
| E9.2 | proposed pose-only | pose | 79.75% | language-aware model |
| E9.3 | proposed RGB/image-only | RGB/image | 57.18% | visual-only model |
| E9.4 | proposed RGB+pose/image | multimodal | 69.87% | first multimodal model |
| E9.5 | research-v2 pose-only | pose | 80.26% | best overall focused model |
| E9.6 | research-v2 RGB/image-only | RGB/image | 57.58% | best visual-only model |
| E9.7 | research-v2 RGB+pose/image | multimodal | 71.82% | best multimodal model |
## Methodology
The focused benchmark uses KSL, CASL, and NSL streams:
- CASL signer-independent word-level pose landmarks.
- KSL word-level pose landmarks.
- NSL image-level pose landmarks.
- CASL signer-independent RGB video frames.
- KSLC/KSL static RGB images.
The proposed encoder uses shared modality encoders with language/task
conditioning, task-specific heads, balanced sampling, supervised contrastive
learning, and metric-learning heads. The research-v2 models add motion-aware
pose features, multi-scale temporal modeling, ArcFace/CosFace-style heads, and
center-loss regularization.
## Loading A Checkpoint
```python
import torch
ckpt = torch.load("models/e9_5_research_v2_pose/pytorch_model.pt", map_location="cpu")
print(ckpt.keys())
```
The checkpoints store the model state dict, normalization statistics, task
metadata, language mappings, and training arguments. Use the code in
`code/experiments/` or the GitHub repository to rebuild the architecture before
loading the state dict.
## Project
GitHub: https://github.com/Yen-hub/AfriSignEncoder
Model repo: https://huggingface.co/luciayen/afrisign-encoder-best-models