GPT2VL (Stackformer V2)

Gated sparse cross-attention Vision-Language model: a frozen GPT-2 small backbone with a Perceiver-style visual resampler and gated cross-attention blocks (spliced before layers 3, 6, 9), trained on Trickxter/COCO2017-captions.

Important: model_trainable.safetensors contains ONLY the trained weights (resampler.* and cross_blocks.*). It does NOT include the frozen GPT-2 backbone or ViT vision encoder weights. To use this model you need:

  1. The base gpt2 weights (loaded via GPT2LMHeadModel.from_pretrained("gpt2"))
  2. The stackformer package / this repo's GPT2VL model definition
  3. This checkpoint's config.json for architecture hyperparameters
  4. model_trainable.safetensors loaded with strict=False on top of the assembled model
Downloads last month
35
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using gurumurthy3/gpt2vl-stackformer-v2 1