Running on Zero Agents GPT2VL Stackformer V2 π A lightweight vision-language model for image captioning