Instructions to use LionelLow/SignSparK_Retrieval with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use LionelLow/SignSparK_Retrieval with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("translation", model="LionelLow/SignSparK_Retrieval")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("LionelLow/SignSparK_Retrieval", device_map="auto") - Notebooks
- Google Colab
- Kaggle
SignSparK-Retrieval
Text-to-gloss models and isolated-sign annotations for the retrieval pipeline of SignSparK: text โ glosses โ retrieved signs โ stitched SignSparK LMDB โ SignSparK fills the transitions.
Code: JianHe0628/SignSparK_Retrieval. Pose data: LionelLow/SignSparK_data.
Contents
| Path | What |
|---|---|
t2g/<dataset>/ |
fine-tuned text-to-gloss model (Hugging Face MBart export) |
annotations/<prefix>.{train,dev,test} |
isolated-sign spans, used to build the sign dictionaries |
annotations/phoenix_iso_signers.tsv |
PHOENIX-2014T video โ signer, for signer-aware retrieval |
predictions/<dataset>/{dev,test}_pres.txt |
the paper's original T2G outputs |
base/<dataset>/ |
vocabulary-trimmed mbart-large-cc25, only needed to train T2G |
<dataset> is CSL-Daily or PHOENIX-2014T; <prefix> is csl_iso or phoenix_iso.
Text-to-gloss
Gloss-level sacreBLEU and WER (%) against the raw LMDB glosses.
| Model | dev BLEU-4 | dev WER | test BLEU-4 | test WER |
|---|---|---|---|---|
t2g/PHOENIX-2014T |
24.53 | 55.02 | 21.01 | 58.92 |
t2g/CSL-Daily |
30.05 | 49.18 | 29.48 | 49.49 |
Our released T2G models have been retrained, but score similarly to the original paper's values. Note that our reproduced retrieval and SignSparK results are in the code README.
Annotations
annotations/<prefix>.<split> is a pickled list with one dict per sign occurrence:
{"video_file", "label", "start", "end", "seq_len"}. start/end are frame indices into
the SignSparK LMDB clip of video_file (end exclusive).
License
The models are fine-tuned from mbart-large-cc25 (vocabulary-trimmed as in GFSLT-VLP). The models, annotations and predictions are derived from PHOENIX-2014T and CSL-Daily and are provided for non-commercial research use only, under the terms of those datasets.
Citation
@inproceedings{low2026signspark,
title={SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning},
author={Low, Jianhe and Symeonidis-Herzig, Alexandre and Ivashechkin, Maksym and Sincan, Ozge Mercanoglu and Bowden, Richard},
booktitle={European Conference on Computer Vision},
pages={648--670},
year={2026},
organization={Springer}
}
Model tree for LionelLow/SignSparK_Retrieval
Base model
facebook/mbart-large-cc25