SignSparK-Retrieval

Text-to-gloss models and isolated-sign annotations for the retrieval pipeline of SignSparK: text โ†’ glosses โ†’ retrieved signs โ†’ stitched SignSparK LMDB โ†’ SignSparK fills the transitions.

Code: JianHe0628/SignSparK_Retrieval. Pose data: LionelLow/SignSparK_data.

Contents

Path What
t2g/<dataset>/ fine-tuned text-to-gloss model (Hugging Face MBart export)
annotations/<prefix>.{train,dev,test} isolated-sign spans, used to build the sign dictionaries
annotations/phoenix_iso_signers.tsv PHOENIX-2014T video โ†’ signer, for signer-aware retrieval
predictions/<dataset>/{dev,test}_pres.txt the paper's original T2G outputs
base/<dataset>/ vocabulary-trimmed mbart-large-cc25, only needed to train T2G

<dataset> is CSL-Daily or PHOENIX-2014T; <prefix> is csl_iso or phoenix_iso.

Text-to-gloss

Gloss-level sacreBLEU and WER (%) against the raw LMDB glosses.

Model dev BLEU-4 dev WER test BLEU-4 test WER
t2g/PHOENIX-2014T 24.53 55.02 21.01 58.92
t2g/CSL-Daily 30.05 49.18 29.48 49.49

Our released T2G models have been retrained, but score similarly to the original paper's values. Note that our reproduced retrieval and SignSparK results are in the code README.

Annotations

annotations/<prefix>.<split> is a pickled list with one dict per sign occurrence: {"video_file", "label", "start", "end", "seq_len"}. start/end are frame indices into the SignSparK LMDB clip of video_file (end exclusive).

License

The models are fine-tuned from mbart-large-cc25 (vocabulary-trimmed as in GFSLT-VLP). The models, annotations and predictions are derived from PHOENIX-2014T and CSL-Daily and are provided for non-commercial research use only, under the terms of those datasets.

Citation

@inproceedings{low2026signspark,
  title={SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning},
  author={Low, Jianhe and Symeonidis-Herzig, Alexandre and Ivashechkin, Maksym and Sincan, Ozge Mercanoglu and Bowden, Richard},
  booktitle={European Conference on Computer Vision},
  pages={648--670},
  year={2026},
  organization={Springer}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for LionelLow/SignSparK_Retrieval

Finetuned
(31)
this model

Dataset used to train LionelLow/SignSparK_Retrieval