dino-vitb16 / README.md
IMvision12's picture
Fix Collection badge link to the current zeromodels collection slug
86293d8 verified
|
Raw
History Blame Contribute Delete
3.62 kB
metadata
pipeline_tag: image-feature-extraction
license: apache-2.0
base_model: facebook/dino-vitb16
library_name: zeromodels
tags:
  - keras
  - zeromodels
  - dino
  - feature-extraction
  - vision
  - arxiv:2104.14294
  - pytorch
  - jax
  - tf

See our collection for all versions of DINO.

Run DINO with Keras 3: JAX, PyTorch, or TensorFlow

GitHub Docs Collection

zeromodels/dino-vitb16

Paper: Emerging Properties in Self-Supervised Vision Transformers (arXiv:2104.14294) · HF Papers

DINO is self-supervised: a student and teacher match across crops of the same image with no labels. The resulting features are semantic for free. These checkpoints are backbones that return tokens / feature maps.

For more details on the model, please go to the upstream model card.

Pure-Keras 3 conversion of facebook/dino-vitb16 for zeromodels. One implementation runs unmodified on TensorFlow / Torch / JAX.

This is a self-supervised backbone (DinoViTModel), not a task head.

✨ Quick start

import os
os.environ["KERAS_BACKEND"] = "torch"  # or "jax" / "tensorflow"

from zeromodels.models.dino import DinoViTModel, DinoImageProcessor

# The processor resizes + ImageNet-normalizes, so build the model with
# include_normalization=False (it would otherwise normalize a second time).
model = DinoViTModel.from_weights(
    "zeromodels/dino-vitb16", include_normalization=False
)
processor = DinoImageProcessor.from_weights("zeromodels/dino-vitb16")

pixel_values = processor("your_image.jpg")["pixel_values"]
features = model(pixel_values, training=False)
print(pixel_values.shape, features.shape)

Load any DINO variant the same way with from_weights("zeromodels/<variant>"):

Variant Hub Backbone
dino-vits16 zeromodels/dino-vits16 ViT-S/16
dino-vits8 zeromodels/dino-vits8 ViT-S/8
dino-vitb16 zeromodels/dino-vitb16 ViT-B/16
dino-vitb8 zeromodels/dino-vitb8 ViT-B/8
dino-resnet50 zeromodels/dino-resnet50 ResNet-50

Tips

  • Set KERAS_BACKEND before importing Keras / zeromodels.
  • The processor normalizes; pair it with include_normalization=False. To skip it, feed raw [0, 255] pixels and keep the default include_normalization=True.
  • dino-resnet50 was converted from torch.hub facebookresearch/dino.
  • See DINO docs and Loading Weights.
  • Community / upstream weights: DinoViTModel.from_weights("hf:facebook/dino-vitb16").

Special Thanks

A huge thank you to the Facebook AI Research DINO authors for creating and releasing these models.

License: Apache 2.0.