--- tags: - image-feature-extraction - timm - transformers pipeline_tag: image-feature-extraction library_name: timm license: apache-2.0 --- # Model card for vit_giant_patch16_lingbot.robbyant A LingBot-Vision ViT-Giant/16 image feature encoder. Pretrained with masked boundary modeling by the paper authors and converted to timm's Eva/DINOv3 implementation. ## Model Notes * Token layout: CLS at index 0, four register tokens at indices 1-4, patch tokens thereafter. The pretrained cfg uses `global_pool='avg'` over patch tokens; pass `global_pool='token'` at creation to reproduce the upstream CLS representation. * fp32 forward outputs match the reference implementation with max abs diff 0.0e+00 on CLS, register, and patch tokens at 512x512 and 384x512. Conversion provenance, the pinned source revision, and per-partition parity metrics are recorded in `manifest.json`. * Converted from https://huggingface.co/robbyant/lingbot-vision-vit-giant at revision `f87d0865a0ae`. The `vit_giant_patch16_lingbot.robbyant` architecture is pending in timm (PR); its pretrained cfg resolves the weights from this repo, so the usage below works on a timm checkout that includes the LingBot entrypoints. ## Model Details - **Model Type:** Image Feature Encoder - **Model Stats:** - Params (M): 1134.5 - GMACs: 1167.15 - Activations (M): 654.86 - Image size: 512 x 512 - **Original:** https://github.com/robbyant/lingbot-vision - **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0) - **Pretrain Dataset:** 161M-image curated web corpus (see paper) - **Papers:** - Vision Pretraining for Dense Spatial Perception: https://arxiv.org/abs/2607.05247 - PyTorch Image Models: https://github.com/huggingface/pytorch-image-models ## Model Usage ### Image Classification ```python from urllib.request import urlopen from PIL import Image import timm import torch img = Image.open(urlopen( 'https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/beignets-task-guide.png' )) model = timm.create_model('vit_giant_patch16_lingbot.robbyant', pretrained=True) model = model.eval() # get model specific transforms (normalization, resize) data_config = timm.data.resolve_model_data_config(model) transforms = timm.data.create_transform(**data_config, is_training=False) output = model(transforms(img).unsqueeze(0)) # unsqueeze single image into batch of 1 top5_probabilities, top5_class_indices = torch.topk(output.softmax(dim=1) * 100, k=5) ``` ### Feature Map Extraction ```python from urllib.request import urlopen from PIL import Image import timm img = Image.open(urlopen( 'https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/beignets-task-guide.png' )) model = timm.create_model( 'vit_giant_patch16_lingbot.robbyant', pretrained=True, features_only=True, ) model = model.eval() # get model specific transforms (normalization, resize) data_config = timm.data.resolve_model_data_config(model) transforms = timm.data.create_transform(**data_config, is_training=False) output = model(transforms(img).unsqueeze(0)) # unsqueeze single image into batch of 1 for o in output: # print shape of each feature map in output print(o.shape) ``` ### Image Embeddings ```python from urllib.request import urlopen from PIL import Image import timm img = Image.open(urlopen( 'https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/beignets-task-guide.png' )) model = timm.create_model( 'vit_giant_patch16_lingbot.robbyant', pretrained=True, num_classes=0, # remove classifier nn.Linear ) model = model.eval() # get model specific transforms (normalization, resize) data_config = timm.data.resolve_model_data_config(model) transforms = timm.data.create_transform(**data_config, is_training=False) output = model(transforms(img).unsqueeze(0)) # output is (batch_size, num_features) shaped tensor # or equivalently (without needing to set num_classes=0) output = model.forward_features(transforms(img).unsqueeze(0)) # output is unpooled, a (1, 1029, 1536) shaped tensor output = model.forward_head(output, pre_logits=True) # output is a (1, num_features) shaped tensor ``` ## Model Comparison Explore the dataset and runtime metrics of this model in timm [model results](https://github.com/huggingface/pytorch-image-models/tree/main/results). ## Citation ```bibtex @article{lingbot-vision2026, title={Vision Pretraining for Dense Spatial Perception}, author={Fu, Zelin and Tan, Bin and Sun, Changjiang and Liu, Shaohui and Zheng, Kecheng and Xu, Yinghao and Zhu, Xing and Shen, Yujun and Xue, Nan}, journal={arXiv preprint arXiv:2607.05247}, year={2026} } ``` ```bibtex @misc{rw2019timm, author = {Ross Wightman}, title = {PyTorch Image Models}, year = {2019}, publisher = {GitHub}, journal = {GitHub repository}, doi = {10.5281/zenodo.4414861}, howpublished = {\url{https://github.com/huggingface/pytorch-image-models}} } ```