--- library_name: skateformer license: mit pipeline_tag: video-classification tags: - skeleton-action-recognition - action-recognition - transformer - eccv2024 - pytorch --- # SkateFormer Official checkpoints for [SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition](https://huggingface.co/papers/2403.09508) (ECCV 2024). SkateFormer partitions joints and frames into four *Skate-Types* (neighbouring/distant joints × local/global frames) and applies **Skate-MSA** within each partition, so attention is spent on the skeletal-temporal relations that actually matter for an action instead of on all joint-frame pairs. At 2.0-3.6M parameters it reaches state-of-the-art accuracy on NTU RGB+D, NTU RGB+D 120, NTU-Inter and NW-UCLA. - 📄 Paper: https://huggingface.co/papers/2403.09508 - 🏠 Project page: https://kaist-viclab.github.io/SkateFormer_site/ - 💻 Code: https://github.com/KAIST-VICLab/SkateFormer ## Checkpoints All 14 released checkpoints live in this one repository, one directory each. Select one with the `subfolder` argument. | `subfolder` | Dataset | Protocol | Modality | Classes | Params | Top-1 | |---|---|---|---|---|---|---| | `ntu60-xsub-joint` | NTU RGB+D 60 | X-Sub | joint | 60 | 3.62M | **92.6%**
J+B: 93.0% | | `ntu60-xsub-bone` | NTU RGB+D 60 | X-Sub | bone | 60 | 3.62M | **92.1%**
J+B: 93.0% | | `ntu60-xview-joint` | NTU RGB+D 60 | X-View | joint | 60 | 3.62M | **97.0%**
J+B: 97.4% | | `ntu60-xview-bone` | NTU RGB+D 60 | X-View | bone | 60 | 3.62M | **96.5%**
J+B: 97.4% | | `ntu120-xsub-joint` | NTU RGB+D 120 | X-Sub | joint | 120 | 3.63M | **87.7%**
J+B: 89.4% | | `ntu120-xsub-bone` | NTU RGB+D 120 | X-Sub | bone | 120 | 3.63M | **88.2%**
J+B: 89.4% | | `ntu120-xset-joint` | NTU RGB+D 120 | X-Set | joint | 120 | 3.63M | **89.3%**
J+B: 91.0% | | `ntu120-xset-bone` | NTU RGB+D 120 | X-Set | bone | 120 | 3.63M | **89.8%**
J+B: 91.0% | | `ntu60-inter-xsub-joint` | NTU-Inter | X-Sub | joint | 11 | 3.61M | **97.1%** | | `ntu60-inter-xview-joint` | NTU-Inter | X-View | joint | 11 | 3.61M | **99.3%** | | `ntu120-inter-xsub-joint` | NTU-Inter 120 | X-Sub | joint | 26 | 3.61M | **92.3%** | | `ntu120-inter-xset-joint` | NTU-Inter 120 | X-Set | joint | 26 | 3.61M | **93.2%** | | `nwucla-joint` | NW-UCLA | official split | joint | 10 | 1.93M | 98.3% | | `nwucla-bone` | NW-UCLA | official split | bone | 10 | 1.93M | 98.3% | Accuracy is the top-1 figure reported in the paper. `J+B` is the E2 ensemble — average the softmax outputs of the joint and bone checkpoints of the same row pair. The paper reports a single NW-UCLA figure with no per-modality or per-ensemble breakdown, so it is listed once rather than attributed to either stream. ## Usage ```bash pip install torch timm huggingface_hub safetensors pip install git+https://github.com/KAIST-VICLab/SkateFormer.git ``` ```python import torch from skateformer import SkateFormer model = SkateFormer.from_pretrained( "JeonghyeokDo/SkateFormer", subfolder="ntu60-xsub-joint" ).eval() # [B, C, T, V, M] — already joint-partitioned, see below x = torch.randn(1, 3, 64, 24, 2) with torch.no_grad(): logits = model(x) # -> [1, 60] print(model.id2label[logits.argmax(-1).item()]) ``` ### Preprocessing SkateFormer does not take raw skeletons directly. A sequence must be (1) converted to the requested modality, (2) resampled to a 64-frame clip, and (3) reordered into skeletal partitions. `skateformer.preprocessing` reproduces the evaluation-time path of the original feeders: ```python import numpy as np from skateformer.preprocessing import prepare_input raw = np.random.randn(3, 300, 25, 2) # [C, T, V, M] raw NTU skeleton (25 joints, 2 people) x, index_t = prepare_input(raw, valid_frame_num=120, layout="ntu", modality="j") with torch.no_grad(): logits = model(x, index_t) ``` Use `layout="nw_ucla"` (20 joints, 1 person) for the NW-UCLA checkpoints, and `modality="b"` for the bone ones — the modality must match the checkpoint you loaded. `index_t` carries the normalised timestamps of the sampled frames (in `[-1, 1]`) and drives the model's temporal index embedding. If omitted, the model assumes a clip that uniformly spans the whole sequence. ### Two-stream ensemble ```python joint = SkateFormer.from_pretrained("JeonghyeokDo/SkateFormer", subfolder="ntu60-xsub-joint").eval() bone = SkateFormer.from_pretrained("JeonghyeokDo/SkateFormer", subfolder="ntu60-xsub-bone").eval() xj, t = prepare_input(raw, valid_frame_num=120, layout="ntu", modality="j") xb, _ = prepare_input(raw, valid_frame_num=120, layout="ntu", modality="b") with torch.no_grad(): probs = (joint(xj, t).softmax(-1) + bone(xb, t).softmax(-1)) / 2 ``` ## Classes `model.id2label` maps class ids to names for the loaded checkpoint. The label sets are NTU RGB+D 60 (60), NTU RGB+D 120 (120), NTU-Inter (11: A50-A60), NTU-Inter 120 (26: A50-A60 + A106-A120) and NW-UCLA (10): 0. pick up with one hand 1. pick up with two hands 2. drop trash 3. walk around 4. sit down 5. stand up 6. donning 7. doffing 8. throw 9. carry ## Limitations - Each checkpoint is trained on one dataset; accuracy drops on skeletons from a different pose estimator or joint layout without fine-tuning. - The joint partition order is part of the model: feeding un-partitioned joints silently produces wrong predictions. ## License MIT, as in the [original repository](https://github.com/KAIST-VICLab/SkateFormer/blob/main/LICENSE). The underlying datasets (NTU RGB+D, NTU RGB+D 120, NW-UCLA) carry their own terms. ## Citation ```bibtex @inproceedings{do2024skateformer, title = {SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition}, author = {Do, Jeonghyeok and Kim, Munchurl}, booktitle = {European Conference on Computer Vision (ECCV)}, year = {2024} } ```