Finetuning midddle layers (vision module layers = 22 to 26 +projector + language layers = 0 to 4) HyperParameters of this model lr: 2e-5 epochs: 10 batch size: 1 grad acumulation step: 4