dino-vitb16 / README.md
IMvision12's picture
Fix Collection badge link to the current zeromodels collection slug
86293d8 verified
|
Raw
History Blame Contribute Delete
3.62 kB
---
pipeline_tag: image-feature-extraction
license: apache-2.0
base_model: facebook/dino-vitb16
library_name: zeromodels
tags:
- keras
- zeromodels
- dino
- feature-extraction
- vision
- arxiv:2104.14294
- pytorch
- jax
- tf
---
## ***See [our collection](https://huggingface.co/collections/zeromodels/dino-v1-v2-v3-6a8eaf5a43e1a5079d6cc817) for all versions of DINO.***
# Run DINO with Keras 3: JAX, PyTorch, or TensorFlow
[![GitHub](https://img.shields.io/badge/GitHub-ZeroModels-black?logo=github)](https://github.com/IMvision12/ZeroModels) [![Docs](https://img.shields.io/badge/Docs-DINO-blue)](https://imvision12.github.io/ZeroModels/dino/) [![Collection](https://img.shields.io/badge/HF-DINO%20collection-yellow)](https://huggingface.co/collections/zeromodels/dino-v1-v2-v3-6a8eaf5a43e1a5079d6cc817)
# zeromodels/dino-vitb16
Paper: [Emerging Properties in Self-Supervised Vision Transformers (arXiv:2104.14294)](https://arxiv.org/abs/2104.14294) · [HF Papers](https://huggingface.co/papers/2104.14294)
DINO is self-supervised: a student and teacher match across crops of the same image with no labels. The resulting features are semantic for free. These checkpoints are backbones that return tokens / feature maps.
For more details on the model, please go to the upstream [model card](https://huggingface.co/facebook/dino-vitb16).
Pure-**Keras 3** conversion of [`facebook/dino-vitb16`](https://huggingface.co/facebook/dino-vitb16) for [zeromodels](https://github.com/IMvision12/ZeroModels). One implementation runs unmodified on **TensorFlow / Torch / JAX**.
This is a **self-supervised backbone** (`DinoViTModel`), not a task head.
## ✨ Quick start
```python
import os
os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow"
from zeromodels.models.dino import DinoViTModel, DinoImageProcessor
# The processor resizes + ImageNet-normalizes, so build the model with
# include_normalization=False (it would otherwise normalize a second time).
model = DinoViTModel.from_weights(
"zeromodels/dino-vitb16", include_normalization=False
)
processor = DinoImageProcessor.from_weights("zeromodels/dino-vitb16")
pixel_values = processor("your_image.jpg")["pixel_values"]
features = model(pixel_values, training=False)
print(pixel_values.shape, features.shape)
```
Load any DINO variant the same way with `from_weights("zeromodels/<variant>")`:
| Variant | Hub | Backbone |
|---|---|---|
| `dino-vits16` | [`zeromodels/dino-vits16`](https://huggingface.co/zeromodels/dino-vits16) | ViT-S/16 |
| `dino-vits8` | [`zeromodels/dino-vits8`](https://huggingface.co/zeromodels/dino-vits8) | ViT-S/8 |
| `dino-vitb16` | [`zeromodels/dino-vitb16`](https://huggingface.co/zeromodels/dino-vitb16) | ViT-B/16 |
| `dino-vitb8` | [`zeromodels/dino-vitb8`](https://huggingface.co/zeromodels/dino-vitb8) | ViT-B/8 |
| `dino-resnet50` | [`zeromodels/dino-resnet50`](https://huggingface.co/zeromodels/dino-resnet50) | ResNet-50 |
## Tips
- Set `KERAS_BACKEND` **before** importing Keras / zeromodels.
- The processor normalizes; pair it with `include_normalization=False`. To skip it, feed raw `[0, 255]` pixels and keep the default `include_normalization=True`.
- `dino-resnet50` was converted from torch.hub `facebookresearch/dino`.
- See [DINO docs](https://imvision12.github.io/ZeroModels/dino/) and [Loading Weights](https://imvision12.github.io/ZeroModels/loading_weights/).
- Community / upstream weights: `DinoViTModel.from_weights("hf:facebook/dino-vitb16")`.
## Special Thanks
A huge thank you to the Facebook AI Research DINO authors for creating and releasing these models.
License: Apache 2.0.