IMvision12's picture
Fix Collection badge link to the current zeromodels collection slug
2b45ba4 verified
|
Raw
History Blame Contribute Delete
5.96 kB
metadata
pipeline_tag: image-classification
license: mit
base_model: timm/swin_base_patch4_window7_224.ms_in1k
library_name: zeromodels
tags:
  - keras
  - zeromodels
  - image-classification
  - swin
  - backbone
  - arxiv:2103.14030
  - pytorch
  - jax
  - tf

See our collection for all versions of Swin Transformer.

Run Swin Transformer with Keras 3: JAX, PyTorch, or TensorFlow

GitHub Docs Collection

zeromodels/swin_base_patch4_window7_224_ms_in1k

Paper: Swin Transformer: Hierarchical Vision Transformer using Shifted Windows (arXiv:2103.14030) · HF Papers

Swin Transformer builds hierarchical feature maps with shifted-window attention. Strong as an ImageNet classifier and as a 4-stage backbone.

For more details on the model, please go to the upstream model card.

Pure-Keras 3 conversion of timm/swin_base_patch4_window7_224.ms_in1k for zeromodels. One implementation runs unmodified on TensorFlow / Torch / JAX.

This is an image-classification / backbone checkpoint (SwinImageClassify / SwinModel).

✨ Quick start

import os
os.environ["KERAS_BACKEND"] = "torch"  # or "jax" / "tensorflow"

from PIL import Image
import numpy as np
from zeromodels.models.swin import SwinImageClassify, SwinModel

model = SwinImageClassify.from_weights("zeromodels/swin_base_patch4_window7_224_ms_in1k")
backbone = SwinModel.from_weights(
    "zeromodels/swin_base_patch4_window7_224_ms_in1k", as_backbone=True
)

image = Image.open("your_image.jpg").convert("RGB")
image = image.resize((224, 224))
x = np.asarray(image, dtype="float32")[None]  # (1, H, W, 3)
print(model(x).shape)  # (1, num_classes)
feats = backbone(x)
print(len(feats), [tuple(f.shape) for f in feats])

Load any Swin Transformer variant the same way with from_weights("zeromodels/<variant>"):

Variant Hub
swin_base_patch4_window12_384_ms_in1k zeromodels/swin_base_patch4_window12_384_ms_in1k
swin_base_patch4_window12_384_ms_in22k zeromodels/swin_base_patch4_window12_384_ms_in22k
swin_base_patch4_window12_384_ms_in22k_ft_in1k zeromodels/swin_base_patch4_window12_384_ms_in22k_ft_in1k
swin_base_patch4_window7_224_ms_in1k zeromodels/swin_base_patch4_window7_224_ms_in1k
swin_base_patch4_window7_224_ms_in22k zeromodels/swin_base_patch4_window7_224_ms_in22k
swin_base_patch4_window7_224_ms_in22k_ft_in1k zeromodels/swin_base_patch4_window7_224_ms_in22k_ft_in1k
swin_large_patch4_window12_384_ms_in22k zeromodels/swin_large_patch4_window12_384_ms_in22k
swin_large_patch4_window12_384_ms_in22k_ft_in1k zeromodels/swin_large_patch4_window12_384_ms_in22k_ft_in1k
swin_large_patch4_window7_224_ms_in22k zeromodels/swin_large_patch4_window7_224_ms_in22k
swin_large_patch4_window7_224_ms_in22k_ft_in1k zeromodels/swin_large_patch4_window7_224_ms_in22k_ft_in1k
swin_small_patch4_window7_224_ms_in1k zeromodels/swin_small_patch4_window7_224_ms_in1k
swin_small_patch4_window7_224_ms_in22k zeromodels/swin_small_patch4_window7_224_ms_in22k
swin_small_patch4_window7_224_ms_in22k_ft_in1k zeromodels/swin_small_patch4_window7_224_ms_in22k_ft_in1k
swin_tiny_patch4_window7_224_ms_in1k zeromodels/swin_tiny_patch4_window7_224_ms_in1k
swin_tiny_patch4_window7_224_ms_in22k zeromodels/swin_tiny_patch4_window7_224_ms_in22k

Tips

  • Set KERAS_BACKEND before importing Keras / zeromodels.
  • SwinImageClassify returns class logits; SwinModel returns features (as_backbone=True for multi-scale stages).
  • See docs and Loading Weights.
  • Upstream / timm checkpoints: SwinImageClassify.from_weights("hf:timm/swin_base_patch4_window7_224.ms_in1k").

Special Thanks

A huge thank you to the Swin Transformer authors and the timm / Hub communities for creating and releasing these models.

License: see YAML license (usually matches the upstream checkpoint).