File size: 5,972 Bytes
8532b73
 
 
 
292f37c
8532b73
b1e7c66
292f37c
b1e7c66
 
f41cc3b
 
b1e7c66
 
 
8532b73
 
1c54f08
8532b73
f41cc3b
8532b73
1c54f08
f41cc3b
292f37c
f41cc3b
 
 
 
 
 
 
292f37c
f41cc3b
 
 
 
8532b73
 
f41cc3b
 
 
 
 
292f37c
f41cc3b
292f37c
f41cc3b
292f37c
f41cc3b
 
 
 
 
 
 
 
8532b73
 
292f37c
f41cc3b
 
 
292f37c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f41cc3b
 
 
292f37c
f41cc3b
292f37c
f41cc3b
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
---
pipeline_tag: image-classification
license: mit
base_model: timm/swin_tiny_patch4_window7_224.ms_in22k
library_name: zeromodels
tags:
- keras
- zeromodels
- image-classification
- swin
- backbone
- arxiv:2103.14030
- pytorch
- jax
- tf
---

## ***See [our collection](https://huggingface.co/collections/zeromodels/swin-transformer-6a8eae9c99ad126043e91850) for all versions of Swin Transformer.***

# Run Swin Transformer with Keras 3: JAX, PyTorch, or TensorFlow

[![GitHub](https://img.shields.io/badge/GitHub-ZeroModels-black?logo=github)](https://github.com/IMvision12/ZeroModels) [![Docs](https://img.shields.io/badge/Docs-Swin-blue)](https://imvision12.github.io/ZeroModels/classification_backbones/) [![Collection](https://img.shields.io/badge/HF-Swin%20collection-yellow)](https://huggingface.co/collections/zeromodels/swin-transformer-6a8eae9c99ad126043e91850)

# zeromodels/swin_tiny_patch4_window7_224_ms_in22k

Paper: [Swin Transformer: Hierarchical Vision Transformer using Shifted Windows (arXiv:2103.14030)](https://arxiv.org/abs/2103.14030) · [HF Papers](https://huggingface.co/papers/2103.14030)

Swin Transformer builds hierarchical feature maps with shifted-window attention. Strong as an ImageNet classifier and as a 4-stage backbone.

For more details on the model, please go to the upstream [model card](https://huggingface.co/timm/swin_tiny_patch4_window7_224.ms_in22k).

Pure-**Keras 3** conversion of [`timm/swin_tiny_patch4_window7_224.ms_in22k`](https://huggingface.co/timm/swin_tiny_patch4_window7_224.ms_in22k) for [zeromodels](https://github.com/IMvision12/ZeroModels). One implementation runs unmodified on **TensorFlow / Torch / JAX**.

This is an **image-classification / backbone** checkpoint (`SwinImageClassify` / `SwinModel`).

## ✨ Quick start

```python
import os
os.environ["KERAS_BACKEND"] = "torch"  # or "jax" / "tensorflow"

from PIL import Image
import numpy as np
from zeromodels.models.swin import SwinImageClassify, SwinModel

model = SwinImageClassify.from_weights("zeromodels/swin_tiny_patch4_window7_224_ms_in22k")
backbone = SwinModel.from_weights(
    "zeromodels/swin_tiny_patch4_window7_224_ms_in22k", as_backbone=True
)

image = Image.open("your_image.jpg").convert("RGB")
image = image.resize((224, 224))
x = np.asarray(image, dtype="float32")[None]  # (1, H, W, 3)
print(model(x).shape)  # (1, num_classes)
feats = backbone(x)
print(len(feats), [tuple(f.shape) for f in feats])
```

Load any Swin Transformer variant the same way with `from_weights("zeromodels/<variant>")`:

| Variant | Hub |
|---|---|
| `swin_base_patch4_window12_384_ms_in1k` | [`zeromodels/swin_base_patch4_window12_384_ms_in1k`](https://huggingface.co/zeromodels/swin_base_patch4_window12_384_ms_in1k) |
| `swin_base_patch4_window12_384_ms_in22k` | [`zeromodels/swin_base_patch4_window12_384_ms_in22k`](https://huggingface.co/zeromodels/swin_base_patch4_window12_384_ms_in22k) |
| `swin_base_patch4_window12_384_ms_in22k_ft_in1k` | [`zeromodels/swin_base_patch4_window12_384_ms_in22k_ft_in1k`](https://huggingface.co/zeromodels/swin_base_patch4_window12_384_ms_in22k_ft_in1k) |
| `swin_base_patch4_window7_224_ms_in1k` | [`zeromodels/swin_base_patch4_window7_224_ms_in1k`](https://huggingface.co/zeromodels/swin_base_patch4_window7_224_ms_in1k) |
| `swin_base_patch4_window7_224_ms_in22k` | [`zeromodels/swin_base_patch4_window7_224_ms_in22k`](https://huggingface.co/zeromodels/swin_base_patch4_window7_224_ms_in22k) |
| `swin_base_patch4_window7_224_ms_in22k_ft_in1k` | [`zeromodels/swin_base_patch4_window7_224_ms_in22k_ft_in1k`](https://huggingface.co/zeromodels/swin_base_patch4_window7_224_ms_in22k_ft_in1k) |
| `swin_large_patch4_window12_384_ms_in22k` | [`zeromodels/swin_large_patch4_window12_384_ms_in22k`](https://huggingface.co/zeromodels/swin_large_patch4_window12_384_ms_in22k) |
| `swin_large_patch4_window12_384_ms_in22k_ft_in1k` | [`zeromodels/swin_large_patch4_window12_384_ms_in22k_ft_in1k`](https://huggingface.co/zeromodels/swin_large_patch4_window12_384_ms_in22k_ft_in1k) |
| `swin_large_patch4_window7_224_ms_in22k` | [`zeromodels/swin_large_patch4_window7_224_ms_in22k`](https://huggingface.co/zeromodels/swin_large_patch4_window7_224_ms_in22k) |
| `swin_large_patch4_window7_224_ms_in22k_ft_in1k` | [`zeromodels/swin_large_patch4_window7_224_ms_in22k_ft_in1k`](https://huggingface.co/zeromodels/swin_large_patch4_window7_224_ms_in22k_ft_in1k) |
| `swin_small_patch4_window7_224_ms_in1k` | [`zeromodels/swin_small_patch4_window7_224_ms_in1k`](https://huggingface.co/zeromodels/swin_small_patch4_window7_224_ms_in1k) |
| `swin_small_patch4_window7_224_ms_in22k` | [`zeromodels/swin_small_patch4_window7_224_ms_in22k`](https://huggingface.co/zeromodels/swin_small_patch4_window7_224_ms_in22k) |
| `swin_small_patch4_window7_224_ms_in22k_ft_in1k` | [`zeromodels/swin_small_patch4_window7_224_ms_in22k_ft_in1k`](https://huggingface.co/zeromodels/swin_small_patch4_window7_224_ms_in22k_ft_in1k) |
| `swin_tiny_patch4_window7_224_ms_in1k` | [`zeromodels/swin_tiny_patch4_window7_224_ms_in1k`](https://huggingface.co/zeromodels/swin_tiny_patch4_window7_224_ms_in1k) |
| `swin_tiny_patch4_window7_224_ms_in22k` | [`zeromodels/swin_tiny_patch4_window7_224_ms_in22k`](https://huggingface.co/zeromodels/swin_tiny_patch4_window7_224_ms_in22k) |

## Tips

- Set `KERAS_BACKEND` **before** importing Keras / zeromodels.
- `SwinImageClassify` returns class logits; `SwinModel` returns features (`as_backbone=True` for multi-scale stages).
- See [docs](https://imvision12.github.io/ZeroModels/classification_backbones/) and [Loading Weights](https://imvision12.github.io/ZeroModels/loading_weights/).
- Upstream / timm checkpoints: `SwinImageClassify.from_weights("hf:timm/swin_tiny_patch4_window7_224.ms_in22k")`.

## Special Thanks

A huge thank you to the Swin Transformer authors and the timm / Hub communities for creating and releasing these models.

License: see YAML `license` (usually matches the upstream checkpoint).