See our collection for all versions of MobileViT.

Run MobileViT with Keras 3: JAX, PyTorch, or TensorFlow

GitHub Docs Collection

kerasformers/mobilevit_xxs_cvnets_in1k

Paper: MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer (arXiv:2110.02178) · HF Papers

MobileViT interleaves MobileNetV2 blocks with small transformers for mobile ImageNet classification (256). For Pascal VOC DeepLabV3 segmentation (512), use update_mobilevit_deeplabv3_model_cards.py.

For more details on the model, please go to the upstream model card.

Pure-Keras 3 conversion of timm/mobilevit_xxs.cvnets_in1k for kerasformers. One implementation runs unmodified on TensorFlow / Torch / JAX.

This is an image-classification / backbone checkpoint (MobileViTImageClassify / MobileViTModel).

✨ Quick start

import os
os.environ["KERAS_BACKEND"] = "torch"  # or "jax" / "tensorflow"

from PIL import Image
from kerasformers.models.mobilevit import (
    MobileViTImageClassify,
    MobileViTModel,
    MobileViTImageProcessor,
)

model = MobileViTImageClassify.from_weights("kerasformers/mobilevit_xxs_cvnets_in1k")
processor = MobileViTImageProcessor.from_weights("kerasformers/mobilevit_xxs_cvnets_in1k")

image = Image.open("your_image.jpg").convert("RGB")
logits = model(processor(image)["pixel_values"], training=False)
print(logits.shape)  # (1, num_classes)

backbone = MobileViTModel.from_weights(
    "kerasformers/mobilevit_xxs_cvnets_in1k", as_backbone=True
)
feats = backbone(processor(image)["pixel_values"], training=False)
print(len(feats), [tuple(f.shape) for f in feats])

Load any MobileViT variant the same way with from_weights("kerasformers/<variant>"):

Variant Hub
mobilevit_s_cvnets_in1k kerasformers/mobilevit_s_cvnets_in1k
mobilevit_xs_cvnets_in1k kerasformers/mobilevit_xs_cvnets_in1k
mobilevit_xxs_cvnets_in1k kerasformers/mobilevit_xxs_cvnets_in1k

Tips

  • Set KERAS_BACKEND before importing Keras / kerasformers.
  • MobileViTImageClassify returns class logits; MobileViTModel returns features (as_backbone=True for multi-scale stages).
  • See docs and Loading Weights.
  • Upstream / timm checkpoints: MobileViTImageClassify.from_weights("hf:timm/mobilevit_xxs.cvnets_in1k").

Special Thanks

A huge thank you to the MobileViT authors and the timm / Hub communities for creating and releasing these models.

License: see YAML license (usually matches the upstream checkpoint).

Downloads last month
50
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zeromodels/mobilevit_xxs_cvnets_in1k

Finetuned
(1)
this model

Collection including zeromodels/mobilevit_xxs_cvnets_in1k

Paper for zeromodels/mobilevit_xxs_cvnets_in1k