--- tags: - image-to-text - vision-language - gpt2 - stackformer --- # GPT2VL (Stackformer V2) Gated sparse cross-attention Vision-Language model: a frozen GPT-2 small backbone with a Perceiver-style visual resampler and gated cross-attention blocks (spliced before layers 3, 6, 9), trained on `Trickxter/COCO2017-captions`. **Important:** `model_trainable.safetensors` contains ONLY the trained weights (`resampler.*` and `cross_blocks.*`). It does NOT include the frozen GPT-2 backbone or ViT vision encoder weights. To use this model you need: 1. The base `gpt2` weights (loaded via `GPT2LMHeadModel.from_pretrained("gpt2")`) 2. The `stackformer` package / this repo's `GPT2VL` model definition 3. This checkpoint's `config.json` for architecture hyperparameters 4. `model_trainable.safetensors` loaded with `strict=False` on top of the assembled model